Eval-Driven Development
Error Analysis Comes First
The most common mistake in starting an eval practice is choosing metrics before looking at data. The order that works is the reverse: pull a sample of real traces — enough to see patterns, reviewed properly rather than skimmed — and read them one by one, writing a short free-text note on each failure. Then group those notes into failure categories that emerge from what you saw, not from a list you downloaded. The distribution is almost always surprising and almost always concentrated: a small number of categories account for most failures, and they are rarely the ones the team was arguing about. Only now do you build metrics, and you build them for the biggest categories first. Metrics chosen before error analysis measure what is easy to measure; metrics chosen after it measure what is actually breaking.
- Read traces before writing graders — the taxonomy must come from your data
- Open coding then clustering: free-text notes first, categories second
- Failure distributions are concentrated — fix the largest bucket, not the loudest complaint
- Repeat error analysis periodically; the distribution shifts as you fix things
The Development Loop
Eval-driven development runs the same loop as test-driven development, with rates in place of assertions. Observe a failure in production or analysis. Write it into the eval set as a case with a grader, and confirm the current system fails it — a case that passes before you change anything is measuring nothing. Make the change: prompt, retrieval, tool, model, control flow. Run the full suite, not just the new case, and read the diff both ways — what improved and what regressed. Accept the change only if the aggregate moves the right way and no critical slice degrades. Then ship, watch production, and start again. The discipline that makes this work is running the whole suite every time; the entire value of the practice is catching the regression you were not looking for.
- New case must fail first, or your grader is not testing what you think
- Always run the full suite — the point is the regression you did not anticipate
- Read wins and losses separately; a flat average can hide a large swap of both
- Commit prompt, grader, and cases together so any result is reproducible
Who Owns the Definition of Good
Evals encode a product judgement, so someone with product authority has to own them. The failure pattern is delegating the rubric to whoever is writing the harness, which quietly hands the definition of quality to an engineer optimising for what is easy to grade. The pattern that works is a single accountable domain expert — the clinician, the lawyer, the support lead, the security analyst — who labels cases personally at the start, resolves disputes about what counts as correct, and stays in the loop as the arbiter when the judge and the reviewers disagree. Engineers build the machinery; the expert owns the ground truth. When multiple stakeholders each hold a veto over the definition of correct, rubrics drift into vague committee language that nothing can reliably grade.
- One accountable owner for ground truth beats a committee with a shared document
- Domain experts label the first cases themselves — do not outsource the seed set
- Engineers own graders and harness; the expert owns what correct means
Making It Routine
Practices that depend on virtue decay. Wire the loop into the mechanics of how the team ships: a fast subset on every pull request, the full suite nightly, results posted where the team already looks, and a standing slot where someone reviews a sample of production traces and files new cases. Keep the fast subset genuinely fast, because a suite that takes half an hour stops being run before a change. Track the size and coverage of the eval set as a metric in its own right — a suite that has not grown in two months is either a finished product or an abandoned practice, and it is rarely the former.
- Fast subset on every change, full suite nightly, trends visible by default
- A recurring trace-review slot is what keeps the set connected to reality
- Treat eval-set growth as a health metric for the team, not just the product
Try It Yourself
Everything in this lesson rests on having cases that came from reality rather than from memory. Ten is enough to start, and this is the afternoon that decides whether the rest of the course is theory or work.
Pick one AI feature you ship, are building, or use every day — if you have none, pick an AI product you rely on and treat your own last two weeks of use as the log. Collect ten real inputs from logs, support tickets, or your own recent sessions, and write down what a good answer looks like for each one before you run anything. Use the Eval Cases Sheet at /templates/eval-cases-template.csv so you are not inventing the columns; its starter rows cover a core task, a hard case, a missing-information case, a must-refuse case, and a formatting case. Then run all ten and count how many come back acceptable. Cases written after you have seen the output only describe what you got, which is why the order matters more than the number.
- All ten inputs are copied from real usage — logs, tickets, or your own sessions — none invented at the desk
- Every what-good-looks-like line was written before the system was run on that case
- You can state the result as a count out of ten and name at least one case it failed
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.