Making Evals a Team Habit
Someone Owns the Eval Set
An eval suite with no owner degrades into a suite with no meaning, and that is the most common way an otherwise competent programme fails. Name a person accountable for the case set: what goes in, what comes out, whether the graders still measure what they claim, and whether the mix still resembles what the product does. Ownership is not the same as doing all the work — cases come from everyone, including support, domain experts, and whoever handled the last incident — but somebody has to hold the standard for what a good case looks like, and that role deserves to be as explicit as service ownership is. Give the owner a defined cadence for reviewing the suite as a whole rather than only reacting to failures, and give them the authority to remove cases, which is the part teams never delegate and therefore never do.
- Name an accountable owner for the case set, not just for the test infrastructure
- Cases come from everyone; the standard for a good case comes from one place
- Schedule whole-suite review rather than only failure-driven attention
- The authority to delete cases is what never happens without an owner
When It Gets Updated, and by What Trigger
Eval sets should change on defined triggers rather than whenever somebody remembers. The reliable set of triggers: every production incident, every complaint that reveals a class rather than a one-off, every new feature or capability, every shift in who the users are or what they ask, and a periodic refresh from current traffic to keep the mix representative. The most valuable of these by a distance is the incident trigger, because those cases are free — a real user found the failure for you and it is guaranteed to matter. Make the path from incident to case short enough that it actually gets walked: a direct route from a production trace to a draft case, a target for how quickly that happens, and a check that the new case fails before the fix and passes after. An incident that has not produced a case within the week almost never produces one.
- Trigger on incidents, class-revealing complaints, new features, and traffic drift
- Incident cases are the cheapest and most certainly relevant ones you will get
- Keep the trace-to-case path short and set a target time for walking it
- Confirm the case fails before the fix and passes after, or it protects nothing
What Blocks a Release
Write down what stops a ship before you need it, because a gate defined during an argument is not a gate. The pattern that holds is hard gates on things that are unambiguous and consequential — safety criteria, schema validity, a floor on the critical slice, and any case protecting a past incident — with report-only signals for movements inside the noise band, since gates that fire on noise get disabled within a month and then nothing is gated at all. Put the result where the decision is actually being made, as a readable diff of which cases flipped in each direction with links to the traces, rather than as a single number in a log nobody opens. And define the override explicitly: who can approve shipping past a failing gate, what they must record, and what follow-up that creates. An override path written down in advance is healthier than one improvised late at night.
- Hard-gate safety, schema validity, critical-slice floors, and incident regression cases
- Report-only inside the noise band; gates that cry wolf get switched off
- Show which cases flipped, with trace links, where the reviewer already is
- Write the override path down, including who records what
The Suite Nobody Looks At
The failure mode worth watching for is not an absent eval suite but a present one that has stopped informing anything: it runs, it is green, and no decision has turned on it in months. The symptoms are recognisable. Nobody can name the last regression it caught. Failures get re-run until they pass. Cases exist that no one can explain. The headline number is quoted in reviews while the underlying cases go unread. This happens because the suite got slow, or noisy, or stale, and each of those has a fix — split fast and comprehensive tiers, prune or quarantine cases that fail for reasons unrelated to their purpose, and refresh the mix from current traffic. The honest test is not how many cases you have but whether anything changed because of them. If the answer for this quarter is nothing, the suite is documentation rather than a control.
- The dangerous state is a green suite no decision has depended on in months
- Symptoms: re-running until green, unexplainable cases, headline numbers quoted unread
- Slow, noisy and stale each have a fix — split tiers, prune, refresh from traffic
- The real metric is whether a decision changed because of the suite
Try It Yourself
A gate nobody wrote down is a gate that gets argued about at the worst possible moment. Write yours now, while nothing is on fire.
For an AI feature you ship, are building, or use daily — with none of your own, for a product your team depends on and would one day have to make this call about — write the release gate on one page: which checks block a release outright, which are posted as warnings only, who may override a block, and what they must record when they do. Keep the blocking list to checks that settle an argument rather than start one. Then write the part teams skip: a short paragraph saying exactly what happens the first Friday afternoon the gate fires on a release someone has already promised.
BLOCKS the release: - [safety criterion] - [schema or format validity] - [floor on the slice where failure costs most: __ %] - [every case protecting a past incident] WARNS ONLY (posted in the pull request, does not block): - [aggregate movement inside the noise band] - [judge-score wobble] Override: [who may approve] records [what], which creates [what follow-up] Friday, 4pm, the gate fires and the release was promised. We: [what actually happens]
- Every blocking check can be settled by pointing at a result, with no discussion needed
- At least one check sits on the warn list, and you can say why it is not a blocker
- The Friday paragraph names who decides and what gets written down, not a general intention
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.