家›Whiatepapers›RAGベンチマーク・ホワイトペーパー

Engineering Whitepaper

Retrieval-Augmented Generation for Enterprise Knowledge

A real-world benchmark from the Goalz.Work Knowledge Pipeline: 93.3% Recall@1 across a 556-chunk, 7-document corpus — at a total processing cost of $0.18 USD.

93.3%

Recall@1

$0.18

Total processing cost

556

Chunks / 7 documents

Recall@1 / @3 / @5

93.3%

Mean Reciprocal Rank

0.933

Corpus size

7 docs

Total cost

$0.18

Executive Summary

Real, reproducible results — not simulated benchmarks

Carmatec’s engineering team built and benchmarked a production-oriented RAG pipeline, codenamed the Knowledge Pipeline, to make an organization’s own documents searchable and reasoned over by AI assistants via MCP.

7 docs

556 retrievable text chunks

93.3 %

Recall@1, @3, and @5

0.933

Mean Reciprocal Rank

$0.18

Total processing cost, USD

How It's Built

A self-hosted stack, purpose-built at every layer

Four specialized data stores, connected only through APIs and message queues — never direct database access — so the pipeline evolves independently of the core application.

Queue

RabbitMQ

Bulk ingestion runs entirely through the message queue — one document or ten thousand use the same code path.

ベクター

Qdrant

Purpose-built vector similarity search over embeddings, with metadata filtering.

Graph

Neo4j

Represents documents and sections as nodes, and their relationships as edges.

Store

モンゴDB

Document and chunk storage alongside the core CakePHP application.

Five stages, every document

1

Parsing

Structure extracted deterministically wherever the format allows.

2

Chunking

12.5% overlap between adjacent chunks to avoid context-cliff failures.

3

Enrichment

LLM-generated header and keywords situate each chunk in context.

4

Embedding + storage

Gemini task-type parameters set explicitly for asymmetric optimization.

Benchmarked

Competitive with — or ahead of — published industry figures

14 of 15 queries retrieved their expected source document as the top-ranked result, measured on a real, curated first-party corpus.

Naive RAG (industry, 70–80%)

ウェブデザイナー

75%

Hybrid retrieval (industry, ~91%)

ウェブデザイナー

91%

Carmatec Knowledge Pipeline

ウェブデザイナー

93.3%

Neg.

Honestly reported negative result: an ablation testing Anthropic’s “Contextual Retrieval” technique found no measurable accuracy gain on this corpus — because each of the seven documents is already topically distinct. We expect a measurable benefit on larger, more homogeneous corpora, and plan to re-test then.

From the Build

Engineering lessons — the kind you only learn by shipping

Every one of these was caught by testing against a real system, not assumed away.

01

Gemini’s embedding API silently accepts requests missing the task-type parameter — easy to miss without reading provider docs directly.

02

Qdrant rejects free-form string identifiers outright; point IDs must be a UUID or unsigned integer.

03

A message queue’s dead-letter routing key must preserve the original routing key — a wildcard override doesn’t work as a wildcard at redelivery.

04

Not every real document is well-formed Markdown — one plain-text spec produced zero chunks until a fallback path was added.

05

Test fixtures and real data must never share an identifier namespace — a collision once caused test cleanup to delete real data.

Where It Applies

Industry-agnostic by design

Document categorization is data, not hardcoded logic — the same architecture applies to any organization managing large volumes of unstructured knowledge.

IT Services

Architecture-decision support grounded in past designs and estimates.

Legal Services

Contract clause and precedent search across a firm’s own case history.

健康管理

Clinical protocol and SOP search grounded in documented procedures.

金融サービス

Regulatory compliance and audit-trail knowledge retrieval.

製造業

SOP and equipment-manual retrieval on the shop floor.

HR & People Ops

Policy search that stays accurate as documents change.

教育

Curriculum and course-material retrieval across a content library.

Government

Regulation and policy search with full source traceability.

What's Next

Sequenced, deliberate next steps

1

Reranking and hybrid (keyword + vector) search — the current production-standard baseline for further accuracy gains.

2

A larger, more heterogeneous ingestion corpus, to re-test the contextual-enrichment ablation.

3

Retrieval-time permission enforcement, closing a potential access-control gap identified during design review.

4

OCR ingestion for scanned and image-based documents, already scoped for employee-document use cases.

5

Staging and production infrastructure replication, sequenced deliberately after this proof.

Get the full whitepaper

Every claim backed by a real, reproducible pipeline run and independently verified queries.

About the author

Written by Carmatec's Leadership

Group CEO & Enterprise AI Solutions Architect, Carmatec

Aromal is a technology entrepreneur and enterprise AI architect with more than 25 years of experience across software engineering, cloud infrastructure, datacenter operations and business transformation. As Group CEO, he stays hands-on in designing the enterprise AI platforms, multi-agent systems and knowledge-driven applications he leads Carmatec to build — including the RAG and enterprise knowledge systems this benchmark reports on. His current technical focus spans Retrieval-Augmented Generation, agent orchestration, Model Context Protocol integrations, and AI governance and production readiness.