What Good Looks Like — and Goodhart's Trap
Properties of a Suite Worth Trusting
A good eval suite is discriminative: it separates systems that differ in quality, which means cases that everything passes and cases nothing passes both carry near-zero information and should be pruned or replaced. It is representative, drawn from the real distribution of usage rather than the tidy examples that came to mind. It is fast enough to run constantly, because a suite run monthly is a report and not a tool. It is trusted, meaning the team believes a failure indicates a real problem — one week of chasing false failures destroys that permanently. And it is versioned alongside the code, because a score is meaningless unless you can say exactly which cases, which graders, and which prompts produced it.
- Prune saturated cases — everything passing means no information
- Representative beats comprehensive: match the real input distribution
- A distrusted suite is worse than none, because it costs time and still gets ignored
- Version cases and graders with code; scores without provenance are decoration
Aggregate Scores Hide the Failures That Matter
A single headline number is convenient and routinely misleading. Systems fail unevenly — by input type, by language, by document length, by customer segment, by whether the request is routine or unusual. An aggregate that rises while your highest-value slice falls is a change you would reject if you could see it, and averaging is precisely what stops you seeing it. Report by slice from the beginning, tag every case with the dimensions you care about, and set separate floors on the slices where failure is expensive. The same logic applies to guardrail metrics: a change that improves helpfulness while increasing unsafe outputs or over-refusals is not an improvement, and only a second metric will tell you.
- Tag cases by slice and report per slice — the average is the least informative view
- Set hard floors on critical slices; aggregate gains must not buy them off
- Pair every primary metric with a guardrail metric that must not move
- Track worst-case and tail behaviour, not just the mean
Goodhart's Trap
When a measure becomes a target it stops being a good measure, and eval suites are unusually easy to game — often unintentionally. You tune prompts against the same fifty cases until the system performs beautifully on them and no better in the world, which is overfitting with extra steps. You adopt a judge that rewards thorough-sounding answers, and your product learns to be verbose. You optimise retrieval recall until the context is stuffed with marginally relevant chunks and answers get worse. Each step was a genuine metric improvement. The defences are structural: hold out a test set the team does not iterate against, refresh cases from production regularly, and keep asking the uncomfortable question of whether the metric still tracks the thing users actually value, or has quietly become its own goal.
- Keep a held-out set you never tune against, and rotate cases in from production
- Suspect any metric that improves for several sprints while users report nothing changed
- Judges create incentives — check what stylistic behaviour yours is rewarding
- Periodically re-derive the metric from user value instead of inheriting last quarter's
The North Star
Under all the machinery sits one question: what fraction of real user tasks does this system complete to an acceptable standard, at what cost and what latency? Everything else — retrieval recall, judge scores, tool-selection accuracy, rubric dimensions — is a diagnostic that helps you improve that number or explain why it moved. Diagnostics are indispensable and they are not the goal, so when a diagnostic and the north star disagree, the north star wins and the diagnostic is the thing that needs fixing. Teams that keep this hierarchy explicit avoid the most common late-stage failure in eval practice: a dashboard full of green metrics attached to a product that people have stopped using.
- Primary metric: task completion to a defined standard, reported with cost and latency
- Component metrics are diagnostics — useful, subordinate, and replaceable
- When a proxy and user value diverge, fix the proxy
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.