Evaluating LLMs and Agents
Why Benchmarks Mislead
Public benchmarks are how the field talks about model quality, and they are systematically weaker evidence than they look. Contamination is the structural problem: benchmarks are published text, models train on published text, so a strong score may measure memorisation rather than capability — and the gap between benchmark performance and performance on genuinely novel variants of the same tasks is a recurring, documented finding. Add benchmark saturation at the frontier, narrow task formats that reward test-taking over usefulness, and vendor marketing incentives, and the practitioner's conclusion follows: public benchmarks are a coarse pre-filter for which models to shortlist, nothing more. The only evaluation that predicts whether a model works for your use case is one built from your use case — your tasks, your data, your definition of good. Nobody publishes that benchmark; you have to build it.
- Contamination: training corpora ingest published benchmarks — scores inflate silently
- Leaderboard deltas near the top rarely predict deltas on your actual workload
- Benchmarks measure narrow formats; your product is not a multiple-choice exam
- Use public scores to shortlist, private evals to decide — never the reverse
LLM-as-Judge — Powerful, Biased, Usable
Many qualities you care about — helpfulness, tone, faithfulness to sources — resist string matching, and human review does not scale to production volume. Using a model to grade outputs closes the gap and now underpins most serious eval pipelines. But a judge model is an instrument with known, reproducible biases, and using one unexamined produces confident nonsense. Position bias: in pairwise comparisons, judges favour one position, so every comparison must run both orderings. Verbosity bias: longer answers score higher at equal quality. Self-preference: models rate outputs resembling their own style more favourably. Judges are also poor at open-ended numeric scoring — a rubric with concrete criteria and a small discrete scale beats "rate this 1–10" every time. The discipline that makes judges trustworthy: validate the judge against a human-labelled sample, measure agreement, and re-validate when you change the judge model or the rubric. An unvalidated judge is an unvalidated metric.
- Swap positions in every pairwise comparison; average or discard on disagreement
- Rubric-based grading with few discrete levels beats open numeric scales
- Watch verbosity and self-preference bias — judge with a different family than you generate with where feasible
- Calibrate against human labels before trusting the judge; recalibrate on every judge change
Harness Design and Agent Regression Suites
An eval harness is the machinery that makes quality measurable: a dataset of cases with inputs and expectations, a runner that executes your actual system against them, graders scoring each result, and reporting that tracks scores across versions. Two design principles matter most. Grade with the cheapest sufficient method — exact checks and code-based assertions where possible, judges only where judgement is genuinely required; deterministic graders are free, fast, and never drift. And evaluate the system, not the model: your prompts, tools, and retrieval in the loop, because that is what ships. For agents, the unit of evaluation shifts from the answer to the trajectory, and non-determinism forces statistical treatment — run each scenario multiple times, assert on pass rates, and grade both end states (did the refund get issued, does the code pass tests) and trajectory properties (were forbidden tools avoided, did it stay under budget). Seed the suite from real production failures; synthetic cases check boxes, incidents encode truth.
- Harness anatomy: cases → runner → graders → versioned reports; keep all four in code
- Prefer code-based graders; spend judge tokens only where judgement is required
- Agent evals grade trajectories and outcomes, at N runs per scenario, with environment resets between runs
- Every incident becomes a case: the regression suite is your institutional memory
The Highest-Leverage Investment
Here is the argument for treating evals as the practitioner's highest-leverage investment. Without them, every change to a prompt, model, or tool is a guess evaluated by vibes — and "it seems better on the three examples I tried" is how silent regressions ship. With them, every improvement compounds: you can upgrade models the week they release because a green suite is your permission slip; you can refactor prompts fearlessly; you can quantify whether the expensive model earns its premium on your tasks — a question no leaderboard answers. Task-completion metrics keep the whole exercise honest: the north star is the fraction of real tasks completed to a defined standard, at known cost and latency — not proxy scores that drift from user value. Teams consistently report the same arc: eval infrastructure feels like overhead in week one and becomes the asset they defend hardest by month six. Your evals encode what "good" means for your product — the one artefact no model vendor can ship you, and the moat that survives every model generation.
- No evals means every change is a bet with no scoreboard — regressions ship silently
- A trusted suite converts model releases from migration risk into same-week upgrades
- North-star metric: task completion to standard, at cost and latency — resist proxy drift
- Start small: twenty real cases beat zero; grow the suite from production, not speculation
- Does Your AI Actually Work? is the full treatment — sourcing cases, judge agreement and sample size, CI, and testing in production
See how many test cases you need before a result means anything.
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.