RAG Architecture In Depth
The Full Pipeline — and Where Quality Is Won
Production RAG is a pipeline, and output quality is the product of every stage: ingest and parse documents, chunk them, embed the chunks, index the vectors, retrieve candidates for a query, rerank them, and assemble the context the model finally sees. The generation step gets the blame when answers are wrong, but in practice most failures are retrieval failures — the right passage never reached the prompt. That inversion is the core practitioner insight: a mediocre model with excellent retrieval beats a frontier model fed the wrong chunks. It also means RAG quality is measurable at every stage independently, and teams that treat it as one opaque "does the answer look right?" system are debugging blind. This lesson walks the stages where the leverage actually lives.
- Retrieval failure masquerades as generation failure — diagnose upstream first
- Parsing matters more than expected: tables, PDFs, and layout-heavy documents lose meaning to naive extraction
- "We use RAG" spans naive top-k cosine similarity to multi-stage hybrid retrieval — the label tells you nothing
Chunking: The Unglamorous High-Leverage Decision
Chunking decides what a retrievable unit of knowledge is, and it quietly bounds the whole system's ceiling. Chunks too small lose context — a sentence about "the previous quarter" retrieves without saying which quarter. Chunks too large dilute the embedding: one vector must represent several topics, matching none well. Fixed-size splitting with overlap is the naive baseline; structure-aware chunking (headings, paragraphs, code blocks) respects semantic boundaries and usually beats it. Stronger patterns decouple matching from reading: parent-document retrieval embeds small precise chunks but hands the model the larger enclosing section; contextual enrichment prepends a document summary or an LLM-generated situating sentence to each chunk before embedding, so isolated fragments carry their own context. Chunking deserves the same experimental rigour as model choice — measured, not defaulted.
- Chunk size trades embedding precision against contextual completeness — there is no universal number
- Structure-aware splitting beats fixed-size; never split mid-table or mid-function
- Parent-document pattern: match on small chunks, generate from their larger context
Hybrid Search and Reranking
Dense embeddings and lexical search fail in complementary ways. Embeddings capture paraphrase and meaning but blur exact identifiers — part numbers, error codes, function names, CVE IDs retrieve poorly as semantics. BM25 keyword scoring nails exact terms but misses reworded concepts. Hybrid search runs both and fuses results, typically with reciprocal rank fusion (RRF), and is the production default because real query streams always contain both kinds. The retrieved candidates then go to a reranker — a cross-encoder that reads query and passage together rather than comparing precomputed vectors, far more accurate and far too slow to run over the whole corpus. The standard shape: cast a wide net with fast hybrid retrieval (top 50–100), then let the reranker choose the handful that enter the prompt.
- BM25 catches what embeddings blur: exact codes, names, identifiers — and vice versa
- RRF fuses ranked lists robustly without tuning score scales against each other
- Cross-encoder rerankers are the single highest-leverage upgrade to a naive RAG stack
Measuring Retrieval: Recall@k, nDCG, and Golden Sets
You cannot tune what you don't measure, and RAG offers concrete retrieval metrics that most teams skip. Build a golden set: real user queries paired with the passages that genuinely answer them, judged by humans — even 50–100 examples transform debugging. Recall@k asks: of the relevant passages, how many appeared in the top k retrieved? It is the ceiling metric — if the answer isn't retrieved, no prompt engineering can save the generation. nDCG additionally rewards putting the most relevant results highest, which matters because models weight early context more heavily. Track these per pipeline change: a new chunk size, embedding model, or reranker becomes an A/B measurement instead of a vibe. End-to-end answer evals (faithfulness, relevance — often LLM-judged) sit on top, but retrieval metrics tell you which layer to fix.
- Golden query set first — everything else in RAG evaluation depends on it
- Recall@k bounds the whole system: missed retrieval is unrecoverable downstream
- nDCG measures ranking quality, not just presence — position in context matters
Compare sentences and see which the model considers close.
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.