The AI Learning Hub Journal

Why Models Invent Case Law

Why a model invents case law that reads perfectlyGENERATION ONLY — NOTHING IS LOOKED UPYou ask for supporting authorityThe model continues the textpredicting what a citation looks like hereSLOTS FILLED BY PLAUSIBILITYparty namesshaped rightreporter formshaped rightcourtshaped rightyearshaped rightEvery part is well-formed.No record was ever consulted.An authority that looks correct and is notRETRIEVAL-BACKED — A CORPUS IS QUERIEDYou ask for supporting authorityThe tool searches a real collectionand writes from what it got backSLOTS FILLED FROM RECORDSparty namesretrievedreporter formretrievedcourtretrievedyearretrievedGrounded in documents that exist.The link back can be followed.An authority you can open and readRETRIEVAL LOWERS THE ODDS — IT DOES NOT RETIRE THE DUTYA retrieved authority can still be summarised wrongly, stretched too far, or cited for a proposition it never supported
A fabricated citation is not a lie — it is a well-formed guess at the shape of one, produced by a system that was never looking anything up

The Mechanism, Not the Malfunction

A fabricated citation is not a bug in the ordinary sense. The model is doing exactly what it was built to do: produce text that is probable given the context. Asked for authority supporting a proposition, it generates a sequence that has the statistical shape of a citation supporting that proposition. If a matching real case is strongly represented in training, the output may be that case. If not, the output is still a well-formed citation, because well-formed citations are what the pattern calls for. There is no internal step where the model consults a database and finds nothing. Understanding this matters practically: it means fabrication is not rare or random, it is the default behaviour whenever the requested authority is not readily reproducible from training.

  • Fabrication is the system working as designed, not an occasional defect to be patched
  • There is no lookup step that can fail and return nothing — the model always produces something
  • Risk rises sharply for narrow propositions, recent developments, and minor jurisdictions
  • Asking the model to be careful or to only cite real cases does not change the mechanism

What Fabrication Looks Like in Practice

The failure has recognisable shapes. Wholly invented cases with plausible party names and a citation in the correct reporter format. Real cases with an invented holding attached. Real cases cited for a proposition they do not support, often adjacent to something they do support. Correct case names paired with wrong citations, wrong courts, or wrong years. Quotations that appear verbatim in nothing. Statutory provisions that do not exist, or exist with different content. Composite authorities that blend two real cases into one. Notably, the model will often defend a fabrication when challenged, supply additional invented detail on request, and produce a confident quotation from a judgment that was never written.

  • Shapes include wholly invented cases, real cases with invented holdings, and correct names with wrong citations
  • Misdescription of real authority is more common and harder to catch than pure invention
  • Challenging the model often produces more invented detail rather than a retraction
  • A fluent verbatim quotation is not evidence the passage exists anywhere

Why It Slips Past Experienced Practitioners

It is tempting to think fabricated citations are only filed by people who were not paying attention, and the record does not support that comfort. The output arrives in the exact register of competent legal research, so the ordinary signals of unreliability are absent. It is typically generated at the point of maximum time pressure. It is frequently produced by someone who did not personally run the query — a junior, a contractor, a client who supplied a draft. And the checking step is boring, repetitive, and almost always confirms nothing is wrong, which is precisely the condition under which humans stop performing checks properly. The vulnerability is structural, and treating it as a matter of individual carelessness is what allows it to keep happening.

  • The output arrives in the register of competent research, so the usual warning signals are absent
  • It surfaces at peak deadline pressure and often via someone other than the person who ran the query
  • Checking is a low-yield vigilance task, which is exactly what humans perform worst
  • Framing this as individual carelessness prevents the process fix that would actually work

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.