The AI Learning Hub Journal

Evals as the Unit Test of Probabilistic Systems

Evals are like unit tests — right up until they are notthe analogy earns its keep on process and fails on everything about the assertion itselfDIMENSIONA UNIT TESTAN EVALTHE OUTCOMEwhat a run gives youBinary. It passed or it failed, andthere is nothing in between.Graded. A score, a rubric band, or ajudgement that has degrees in it.REPEATABILITYrun it twiceDeterministic. The same input givesthe same result every single time.Probabilistic. The same input cangive a different answer next run.ONE CASEhow much it tells youOne failing case is conclusive, andit localises a real defect.One case proves almost nothing onits own. You need many, sampled well.THE THRESHOLDwhat counts as passingEvery assertion must hold. A singlefailure fails the whole build.An aggregate bar across a suite, withthe run-to-run spread reported too.WHAT CARRIES OVER FROM TESTINGFast feedback, while the change is still in your headRuns automatically on every change, not when rememberedLives in version control beside the thing it testsCan fail the build and hold back a releaseEvery failure found becomes a permanent case in the suiteWHAT DOES NOT SURVIVE THE MOVEExact-match assertions over free-form outputOne run being enough to call it passed or failedA single case pinpointing a single broken lineA green suite meaning the behaviour is actually correctThresholds chosen without knowing the run-to-run spreadMany cases and an aggregate threshold do the job that one assertion used to doA single run is a sample of the behaviour, so treat a one-run difference as noise until it repeatsKeep the pipeline habits of testing and replace its arithmetic — that is the whole translation
Borrow the discipline of testing and the place it sits in the pipeline — not its idea of what a pass means

From Assertion to Distribution

A unit test asserts that one input yields one output. That contract is unavailable here: sampling makes outputs stochastic, valid answers have many surface forms, and multi-step systems compound small variations into different trajectories. Evals keep the shape of testing while changing the assertion. Instead of asserting equality you assert a property — the answer contains the correct figure, the JSON validates against schema, the refund was issued, no forbidden tool was called. Instead of pass or fail on one run, you assert on a rate across many. The mental shift that unlocks everything: an eval is not a test that passes, it is a measurement with a value, a distribution, and a trend. You gate on thresholds over that measurement the way you gate on latency budgets, not the way you gate on compilation.

  • Assert properties and invariants, never exact strings
  • The result of an eval is a number with uncertainty, not a boolean
  • Thresholds are product decisions — pick them deliberately, then hold the line

The Grading Ladder

Choose the cheapest grader that can actually decide the question, and climb only when forced. Deterministic checks come first: schema validation, exact or fuzzy match against a reference, regex for required identifiers, unit tests over generated code, database state after a tool call. These are free, instant, and never drift. Next come programmatic heuristics and retrieval metrics — recall of the right chunk, correct tool selected, run stayed under budget. Only then do you reach for a model judge, and only for genuinely subjective properties: faithfulness to a source, tone, whether an explanation would satisfy the asker. Most teams invert this ladder, sending everything to a judge because it is one line of code, then discover their metric is an unvalidated model grading another model with no ground truth anywhere in the loop.

  • Deterministic graders: free, fast, drift-proof — use them wherever the answer is checkable
  • Code execution is the strongest grader available when the output is code
  • Reserve judge tokens for properties no assertion can express
  • Every judge you add is another model you now have to validate

Levels of the Suite

Mirror the testing pyramid. Component evals target one unit — a classifier prompt, an extraction step, the retriever alone — with many cases, cheap graders, and fast feedback; when quality drops these tell you which stage broke. Pipeline evals run the assembled system end to end on realistic inputs, catching the interaction bugs no component test sees. Trajectory evals apply to agents, grading the path as well as the destination: tools chosen, steps taken, budget consumed, gates respected. Keep the base of the pyramid wide because it is where debugging happens, but never let it substitute for the top — a system can pass every component eval and still fail users, because the failure lives in how the parts compose.

  • Component evals localise faults; end-to-end evals prove the product works
  • Retrieval quality deserves its own metrics — RAG failures are usually retrieval failures
  • Agent evals grade trajectory and outcome, because a right answer via a forbidden path is a defect

Flakiness Is Data

In conventional CI a flaky test is a defect in the test. In eval suites intermittent failure is often a true report about your system: a case that passes seventy percent of the time is telling you that seventy percent of users asking that question get a correct answer. Do not chase it to green by loosening the grader. Record the rate, decide whether it is acceptable for that slice, and treat movement in it as signal. The genuinely broken cases are the ones where the grader is wrong — ambiguous expected answers, rubrics two reviewers read differently, environments that leak state between runs. Fix grader flakiness ruthlessly and preserve system flakiness faithfully; confusing the two is how suites lose the trust that makes them useful.

  • Intermittent pass rates describe reality — record them rather than suppressing them
  • Grader nondeterminism is a bug; system nondeterminism is a measurement
  • Reset environment state between runs or you are measuring contamination

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.