The AI Learning Hub Journal

What Good Looks Like — and Goodhart's Trap

When the metric keeps rising and the value does nothighlowoptimisation pressure over successive releasesDIVERGENCE POINTthe metric becomes the targetand stops tracking the thingthey agreewhile nobodyis pushingthe metric you reportthe value it stood forHOW THE GAP GETS OPENEDOptimising for the judgethe model learns what the grader rewards —warmth, confidence, agreement — ratherthan what the task actually asked forVerbositycoverage-shaped rubrics pay by the yard,so answers get longer without gettingmore useful to the person reading themBenchmark shapeanswers get fitted to the format of thetest set — its phrasing, its options, itslength — instead of to the real taskStripped hedginguncertainty scores badly, so it is removedfrom answers that genuinely are uncertain,and the score rises as the honesty fallsWHAT KEEPS A PROXY HONESTHold a set backnever optimised against, androtated once it starts to leakPair every metricadd a counter-metric that gamingwould push the wrong wayRead raw outputssample by hand every cycle —scores hide what reading findsRetire saturated onesa number that has stopped movinghas stopped measuring anythingThe divergence is invisible from the dashboard — it only shows up when someone reads the outputs
A proxy is trustworthy only while nobody is pushing on it — the moment it becomes the target, it starts to part company with the thing it stood for

Properties of a Suite Worth Trusting

A good eval suite is discriminative: it separates systems that differ in quality, which means cases that everything passes and cases nothing passes both carry near-zero information and should be pruned or replaced. It is representative, drawn from the real distribution of usage rather than the tidy examples that came to mind. It is fast enough to run constantly, because a suite run monthly is a report and not a tool. It is trusted, meaning the team believes a failure indicates a real problem — one week of chasing false failures destroys that permanently. And it is versioned alongside the code, because a score is meaningless unless you can say exactly which cases, which graders, and which prompts produced it.

  • Prune saturated cases — everything passing means no information
  • Representative beats comprehensive: match the real input distribution
  • A distrusted suite is worse than none, because it costs time and still gets ignored
  • Version cases and graders with code; scores without provenance are decoration

Aggregate Scores Hide the Failures That Matter

A single headline number is convenient and routinely misleading. Systems fail unevenly — by input type, by language, by document length, by customer segment, by whether the request is routine or unusual. An aggregate that rises while your highest-value slice falls is a change you would reject if you could see it, and averaging is precisely what stops you seeing it. Report by slice from the beginning, tag every case with the dimensions you care about, and set separate floors on the slices where failure is expensive. The same logic applies to guardrail metrics: a change that improves helpfulness while increasing unsafe outputs or over-refusals is not an improvement, and only a second metric will tell you.

  • Tag cases by slice and report per slice — the average is the least informative view
  • Set hard floors on critical slices; aggregate gains must not buy them off
  • Pair every primary metric with a guardrail metric that must not move
  • Track worst-case and tail behaviour, not just the mean

Goodhart's Trap

When a measure becomes a target it stops being a good measure, and eval suites are unusually easy to game — often unintentionally. You tune prompts against the same fifty cases until the system performs beautifully on them and no better in the world, which is overfitting with extra steps. You adopt a judge that rewards thorough-sounding answers, and your product learns to be verbose. You optimise retrieval recall until the context is stuffed with marginally relevant chunks and answers get worse. Each step was a genuine metric improvement. The defences are structural: hold out a test set the team does not iterate against, refresh cases from production regularly, and keep asking the uncomfortable question of whether the metric still tracks the thing users actually value, or has quietly become its own goal.

  • Keep a held-out set you never tune against, and rotate cases in from production
  • Suspect any metric that improves for several sprints while users report nothing changed
  • Judges create incentives — check what stylistic behaviour yours is rewarding
  • Periodically re-derive the metric from user value instead of inheriting last quarter's

The North Star

Under all the machinery sits one question: what fraction of real user tasks does this system complete to an acceptable standard, at what cost and what latency? Everything else — retrieval recall, judge scores, tool-selection accuracy, rubric dimensions — is a diagnostic that helps you improve that number or explain why it moved. Diagnostics are indispensable and they are not the goal, so when a diagnostic and the north star disagree, the north star wins and the diagnostic is the thing that needs fixing. Teams that keep this hierarchy explicit avoid the most common late-stage failure in eval practice: a dashboard full of green metrics attached to a product that people have stopped using.

  • Primary metric: task completion to a defined standard, reported with cost and latency
  • Component metrics are diagnostics — useful, subordinate, and replaceable
  • When a proxy and user value diverge, fix the proxy

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.