Testing a System That Answers From Your Documents
Two Failures Wearing One Face
When a document-grounded answer is wrong there are two entirely different causes and one symptom. Either the retrieval step never surfaced the passage containing the answer, in which case no amount of prompt work on the generation step will help, or the right passage was retrieved and the model failed to use it — ignored it, misread it, or blended it with something it already believed. Teams that report a single end-to-end accuracy number cannot tell these apart, and so they spend weeks tuning the wrong stage. Instrument the pipeline so every eval case records what was retrieved, in what order, and what the model was actually given, then grade the two stages separately before grading the whole. The end-to-end number is what the user experiences and it stays the headline; the stage numbers are what tell you where the next sprint should go.
- One wrong answer, two root causes: nothing found, or something found and misused
- Record the retrieved passages and their order on every eval case
- Grade retrieval and generation separately, then report end to end
- A single accuracy number sends teams to tune the stage that was already fine
Did It Find the Right Source at All?
The retrieval question is simple to state: for each case, was a passage that actually contains the answer among the ones the system pulled back? Label each case with its supporting passages once, then measure the share of cases where at least one correct passage appears in the top k results. That is what recall@k means in plain terms, and k should be the number of passages your prompt actually includes rather than the number your search layer returns. Two habits make the measurement useful. Report it at several values of k, because the gap between recall at three and recall at twenty tells you whether the passage is missing from the index entirely or merely ranked too low, and those need completely different fixes. And check where the correct passage lands, since models attend unevenly across a long context and a correct passage in the last position is not reliably as good as one near the front.
- Label the supporting passages per case once and reuse them across every run
- Recall@k: the share of cases where a correct passage sits in the top k you pass in
- Measure at several values of k — missing and mis-ranked are different problems
- Position matters; a correct passage buried deep is not the same as one near the top
Is the Answer Actually Grounded?
Retrieval succeeding does not mean the answer used what was retrieved. Groundedness asks a narrower question than correctness: is every claim in the output supported by the passages that were supplied? Grade it claim by claim rather than as a holistic impression — split the answer into its factual assertions and check each against the provided context, which is a job a judge model does reasonably well when the criterion is that concrete and badly when asked whether an answer seems grounded overall. Track two distinct failures separately. Unsupported claims are content the model added from its own parameters; omissions leave out something the context plainly contained. And keep groundedness apart from correctness, because an answer can be faithfully grounded in a retrieved passage that is itself stale or wrong, which is a corpus problem fixed somewhere else entirely.
- Split the answer into claims and check each one against the supplied context
- Judges grade concrete per-claim support well and holistic groundedness badly
- Separate unsupported additions from omissions of what the context contained
- Grounded and correct are different: a faithful answer to a stale source is still wrong
The Fluent Answer That Cites Nothing
The characteristic failure of these systems is not an obvious error, it is a well-written answer assembled from the model's own knowledge when retrieval returned nothing useful — fluent, plausible, correctly formatted, unsupported. It passes a helpfulness rubric, it reads better than a careful hedged answer, and reviewers skimming a sample will approve it. Build the cases that catch it deliberately. Questions whose answers are genuinely absent from the corpus, where the only correct behaviour is to say so. Questions where the corpus holds a near miss on the topic, which invites confident blending. And questions whose answer changed, so the parametric answer and the corpus answer differ and you can see which one the system actually used. Require citations and then verify them mechanically, because a citation that does not support the sentence attached to it is worse than no citation at all.
- Include unanswerable cases where the only correct output is admitting the gap
- Add near-miss cases that invite blending corpus content with parametric knowledge
- Cases whose answer has changed reveal which source the system really used
- Verify citations mechanically — an unsupporting citation manufactures false confidence
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.