Evals as the Unit Test of Probabilistic Systems
From Assertion to Distribution
A unit test asserts that one input yields one output. That contract is unavailable here: sampling makes outputs stochastic, valid answers have many surface forms, and multi-step systems compound small variations into different trajectories. Evals keep the shape of testing while changing the assertion. Instead of asserting equality you assert a property — the answer contains the correct figure, the JSON validates against schema, the refund was issued, no forbidden tool was called. Instead of pass or fail on one run, you assert on a rate across many. The mental shift that unlocks everything: an eval is not a test that passes, it is a measurement with a value, a distribution, and a trend. You gate on thresholds over that measurement the way you gate on latency budgets, not the way you gate on compilation.
- Assert properties and invariants, never exact strings
- The result of an eval is a number with uncertainty, not a boolean
- Thresholds are product decisions — pick them deliberately, then hold the line
The Grading Ladder
Choose the cheapest grader that can actually decide the question, and climb only when forced. Deterministic checks come first: schema validation, exact or fuzzy match against a reference, regex for required identifiers, unit tests over generated code, database state after a tool call. These are free, instant, and never drift. Next come programmatic heuristics and retrieval metrics — recall of the right chunk, correct tool selected, run stayed under budget. Only then do you reach for a model judge, and only for genuinely subjective properties: faithfulness to a source, tone, whether an explanation would satisfy the asker. Most teams invert this ladder, sending everything to a judge because it is one line of code, then discover their metric is an unvalidated model grading another model with no ground truth anywhere in the loop.
- Deterministic graders: free, fast, drift-proof — use them wherever the answer is checkable
- Code execution is the strongest grader available when the output is code
- Reserve judge tokens for properties no assertion can express
- Every judge you add is another model you now have to validate
Levels of the Suite
Mirror the testing pyramid. Component evals target one unit — a classifier prompt, an extraction step, the retriever alone — with many cases, cheap graders, and fast feedback; when quality drops these tell you which stage broke. Pipeline evals run the assembled system end to end on realistic inputs, catching the interaction bugs no component test sees. Trajectory evals apply to agents, grading the path as well as the destination: tools chosen, steps taken, budget consumed, gates respected. Keep the base of the pyramid wide because it is where debugging happens, but never let it substitute for the top — a system can pass every component eval and still fail users, because the failure lives in how the parts compose.
- Component evals localise faults; end-to-end evals prove the product works
- Retrieval quality deserves its own metrics — RAG failures are usually retrieval failures
- Agent evals grade trajectory and outcome, because a right answer via a forbidden path is a defect
Flakiness Is Data
In conventional CI a flaky test is a defect in the test. In eval suites intermittent failure is often a true report about your system: a case that passes seventy percent of the time is telling you that seventy percent of users asking that question get a correct answer. Do not chase it to green by loosening the grader. Record the rate, decide whether it is acceptable for that slice, and treat movement in it as signal. The genuinely broken cases are the ones where the grader is wrong — ambiguous expected answers, rubrics two reviewers read differently, environments that leak state between runs. Fix grader flakiness ruthlessly and preserve system flakiness faithfully; confusing the two is how suites lose the trust that makes them useful.
- Intermittent pass rates describe reality — record them rather than suppressing them
- Grader nondeterminism is a bug; system nondeterminism is a measurement
- Reset environment state between runs or you are measuring contamination
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.