Hallucination: Mechanisms & Mitigations
The Mechanism: Plausibility Is the Objective
Hallucination is not a defect bolted onto LLMs — it is the training objective showing through. A model trained to predict likely next tokens learns to produce text that is statistically plausible, and truth is only correlated with plausibility, not identical to it. For facts seen thousands of times in training, the most probable continuation is also the true one. For rare facts — a minor person's biography, a niche API's parameters, a specific citation — the model has weak or no memorised signal and interpolates from patterns instead: it produces what such an answer typically looks like. The output is fluent, well-formatted, confidently phrased, and wrong. This is why hallucinations cluster exactly where verification is hardest: specific, rare, detailed claims — and why the failure feels like fabrication rather than error.
- The loss rewards plausible text; factuality is a frequently-correlated side effect, not a constraint
- Fluency and confidence are decoupled from accuracy — style signals nothing about truth
- Citations, URLs, and API signatures are prime hallucination targets: highly structured, easily patterned, rarely memorised
Why Training Makes It Worse Before It Makes It Better
Post-training adds its own pressures. Human raters reward confident, complete, helpful-sounding answers — and unknowingly penalise honest uncertainty, teaching models that a fluent guess beats "I don't know." Sycophancy compounds this: models learn to agree with a user's framing even when it embeds a false premise. Calibration degrades too — base models' token probabilities track their actual accuracy reasonably well, and preference tuning distorts this, producing uniform confidence across correct and incorrect claims. Sampling adds a final layer: at nonzero temperature, a model that would most probably produce a correct fact can still sample an incorrect neighbour. Frontier labs now explicitly train abstention as a first-class behaviour — partial progress, not a solved problem — which is why mitigation remains architectural rather than a patch.
The Mitigation Stack
No single technique eliminates hallucination; production systems layer defences that each catch what the previous missed. Grounding comes first: retrieve relevant documents (RAG) and instruct the model to answer only from them, with citations — shifting the task from recall to reading comprehension, which models do far better. Tool use goes further: let the model call search, databases, or code execution instead of trusting its weights. Constrained output (schemas, enumerated choices) removes room for free-form fabrication in structured tasks. Verification layers — a second model checking claims against sources, or self-consistency across multiple samples — catch residual errors. Finally, human review guards consequential actions. Design to the consequence: a wrong movie recommendation and a wrong drug interaction warrant very different stacks.
- Grounding + citations converts recall into comprehension and makes answers auditable
- Tool calls beat parametric memory for anything current, rare, or precise
- Match mitigation depth to consequence-of-error — a system design decision, not a model setting
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.