Engineering Whitepaper
A real-world benchmark from the Goalz.Work Knowledge Pipeline: 93.3% Recall@1 across a 556-chunk, 7-document corpus — at a total processing cost of $0.18 USD.
Recall@1
Total processing cost
Chunks / 7 documents
Recall@1 / @3 / @5
93.3%
Mean Reciprocal Rank
0.933
Corpus size
7 docs
Total cost
$0.18
Executive Summary
Carmatec’s engineering team built and benchmarked a production-oriented RAG pipeline, codenamed the Knowledge Pipeline, to make an organization’s own documents searchable and reasoned over by AI assistants via MCP.
556 retrievable text chunks
Recall@1, @3, and @5
Mean Reciprocal Rank
Total processing cost, USD
How It's Built
Four specialized data stores, connected only through APIs and message queues — never direct database access — so the pipeline evolves independently of the core application.
Bulk ingestion runs entirely through the message queue — one document or ten thousand use the same code path.
Purpose-built vector similarity search over embeddings, with metadata filtering.
Represents documents and sections as nodes, and their relationships as edges.
Document and chunk storage alongside the core CakePHP application.
Structure extracted deterministically wherever the format allows.
12.5% overlap between adjacent chunks to avoid context-cliff failures.
LLM-generated header and keywords situate each chunk in context.
Gemini task-type parameters set explicitly for asymmetric optimization.
Benchmarked
14 of 15 queries retrieved their expected source document as the top-ranked result, measured on a real, curated first-party corpus.
Honestly reported negative result: an ablation testing Anthropic’s “Contextual Retrieval” technique found no measurable accuracy gain on this corpus — because each of the seven documents is already topically distinct. We expect a measurable benefit on larger, more homogeneous corpora, and plan to re-test then.
From the Build
Every one of these was caught by testing against a real system, not assumed away.
Gemini’s embedding API silently accepts requests missing the task-type parameter — easy to miss without reading provider docs directly.
Qdrant rejects free-form string identifiers outright; point IDs must be a UUID or unsigned integer.
A message queue’s dead-letter routing key must preserve the original routing key — a wildcard override doesn’t work as a wildcard at redelivery.
Not every real document is well-formed Markdown — one plain-text spec produced zero chunks until a fallback path was added.
Test fixtures and real data must never share an identifier namespace — a collision once caused test cleanup to delete real data.
Where It Applies
Document categorization is data, not hardcoded logic — the same architecture applies to any organization managing large volumes of unstructured knowledge.
Architecture-decision support grounded in past designs and estimates.
Contract clause and precedent search across a firm’s own case history.
Clinical protocol and SOP search grounded in documented procedures.
Regulatory compliance and audit-trail knowledge retrieval.
SOP and equipment-manual retrieval on the shop floor.
Policy search that stays accurate as documents change.
Curriculum and course-material retrieval across a content library.
Regulation and policy search with full source traceability.
What's Next
Reranking and hybrid (keyword + vector) search — the current production-standard baseline for further accuracy gains.
A larger, more heterogeneous ingestion corpus, to re-test the contextual-enrichment ablation.
Retrieval-time permission enforcement, closing a potential access-control gap identified during design review.
OCR ingestion for scanned and image-based documents, already scoped for employee-document use cases.
Staging and production infrastructure replication, sequenced deliberately after this proof.
Every claim backed by a real, reproducible pipeline run and independently verified queries.
About the author
Group CEO & Enterprise AI Solutions Architect, Carmatec
Aromal is a technology entrepreneur and enterprise AI architect with more than 25 years of experience across software engineering, cloud infrastructure, datacenter operations and business transformation. As Group CEO, he stays hands-on in designing the enterprise AI platforms, multi-agent systems and knowledge-driven applications he leads Carmatec to build — including the RAG and enterprise knowledge systems this benchmark reports on. His current technical focus spans Retrieval-Augmented Generation, agent orchestration, Model Context Protocol integrations, and AI governance and production readiness.