Retrieval Helps. It Does Not Cure.
How Retrieval-Backed Research Changes the Picture
Research tools built on retrieval work differently from a bare chatbot. Instead of generating an answer from model weights alone, they search an actual corpus of primary and secondary sources, retrieve relevant documents, and instruct the model to answer using only those documents, usually with links back to what was retrieved. This is a genuine improvement and it substantially reduces the wholly-invented-case failure, because the citations come from a real index rather than from the model's learned patterns. If a firm is going to use AI for legal research at all, a properly grounded, source-linked tool is a materially safer starting point than a general assistant. That is a real distinction and worth insisting on in procurement.
- Grounded tools search a real corpus and cite what they retrieved, rather than generating from weights
- This substantially reduces wholly-invented citations — the citation comes from a real index
- A source-linked, grounded tool is materially safer than a general assistant for research
- Insist on visible source links in procurement; a tool without them is not grounded in any useful sense
The Failure Modes Retrieval Does Not Remove
Grounding narrows the failure surface without closing it. The model still summarises what was retrieved and can misstate a holding, overstate how squarely a case supports a point, or lose a crucial qualification. Retrieval can surface a real case that is genuinely on point but has since been overruled, distinguished, or superseded by statute, and currency-checking is a separate function that not every tool performs. Coverage gaps are invisible: if the corpus lacks a jurisdiction or a court level, the tool answers confidently from what it does have. Chunked retrieval can sever a passage from a qualification appearing elsewhere in the judgment. And a generated synthesis can drift from the sources beneath it while still displaying them as citations.
- Misdescription of retrieved authority survives grounding entirely — the summary is still generated
- Overruled, distinguished, or superseded authority can be retrieved and cited as current
- Corpus coverage gaps are silent; the tool answers from what it has without signalling what it lacks
- Chunking can separate a passage from the qualification that changes its meaning
- Measured, not asserted: a Stanford RegLab study of leading AI legal research tools (Journal of Empirical Legal Studies, 2025) found hallucination rates from roughly one in six to one in three of the queries tested
Reading a Grounded Answer Correctly
The right mental model for a grounded research tool is a fast, tireless, and occasionally careless researcher who always hands you the documents. The value is in the documents. The prose summary is a navigational aid — useful for deciding what to read, never a substitute for reading it. In practice this means clicking through to every source before relying on the answer, reading enough of each source to confirm it says what the summary claims, and checking currency independently. It also means noticing when a tool returns a confident answer with thin or tangential sources, which is a signal that the corpus did not contain a good answer and the model synthesised around the gap.
- The retrieved documents are the product; the summary is a navigational aid to them
- Click through to every source and read enough to confirm the characterisation holds
- Check currency separately unless the tool explicitly performs and displays that function
- A confident answer resting on thin or tangential sources means the corpus fell short
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.