Human Review That Scales
Where Judgement Cannot Be Delegated
Automated graders handle anything with a checkable definition of correct — schema validity, a computed figure, a required fact, a forbidden action — and model judges extend that to criteria you can write down explicitly. What remains is genuinely irreducible: whether an answer is appropriate for this customer in this situation, whether the tone fits, whether advice is safe given domain norms a rubric cannot fully capture, and whether the thing is actually useful rather than merely responsive. These are the cases where the rubric you would have to write is the judgement itself. Human review is therefore not a stage that teams graduate out of once they automate properly; it is a permanent tier that shrinks as your rubrics improve and never reaches zero. Treat expert attention as the scarcest input in the whole eval programme and design the process around spending it well.
- Automate anything checkable; use judges for anything you can articulate
- What is left is judgement a rubric cannot express — appropriateness, tone, real usefulness
- Human review is a permanent tier, not a temporary one
- Expert attention is the scarce resource; the process exists to allocate it
Sample, Do Not Review Everything
Reviewing every case is neither possible nor useful, and the instinct to try produces a process that gets quietly abandoned. Sample instead, and sample deliberately rather than uniformly. A small random sample gives an unbiased estimate and keeps you honest about the base rate. On top of that, stratify: over-sample the slices that matter most, the cases automated graders flagged as borderline, the cases where two graders disagreed, and the cases that changed verdict between the last version and this one. Uniform random sampling spends most of its budget confirming that easy cases still pass, which is precisely what automation already told you. Choose the sample size from the precision you need rather than from what feels thorough, and record which sampling scheme produced a number, because comparing a stratified review against last quarter's random one compares nothing at all.
- A small random sample keeps the base rate honest; stratify on top of it
- Over-sample high-stakes slices, borderline cases, and verdict changes
- Uniform sampling re-confirms what automation already established
- Record the sampling scheme with the number or cross-run comparisons are meaningless
Making Two Reviewers Agree
Two competent reviewers given the same output and the same one-line criterion will disagree far more than either expects, and unmeasured disagreement means your human numbers are noise with a serious face on. Measure it: have a portion of every batch labelled independently by two people and track the agreement rate as a standing metric. When it is low, the rubric is the problem rather than the reviewers — the fix is to convert vague adjectives into observable conditions, anchor each point of the scale with worked examples including the boundaries, and write explicit rules for the cases people keep resolving differently. Run a calibration session on a shared set before a batch, discuss the disagreements instead of averaging them away, and reissue the rubric. Agreement between humans is also the ceiling for any model judge validated against those labels, which makes this a prerequisite rather than a refinement.
- Double-label part of every batch and track agreement as a standing metric
- Low agreement is a rubric defect: replace adjectives with observable conditions
- Anchor every scale point with worked examples, especially at the boundaries
- Human agreement caps the accuracy of any judge trained against those labels
Fatigue, Drift, and Attention That Discriminates
Reviewer quality degrades in ways that show up as nothing on a dashboard. Long batches produce fatigue, and labels late in a session are measurably worse than labels early in it. Standards drift over weeks as reviewers recalibrate against whatever they have been seeing, so a score from one quarter and a score from the next are not directly comparable. Presentation order matters too, since whichever output appears first carries an advantage. Counter all three mechanically: cap batch length, rotate a small set of gold cases with known answers through every batch to detect drift, blind and randomise the order of candidates, and periodically re-run an old batch to check that today's reviewers still agree with the earlier labels. Then spend what remains where it discriminates — on the cases where two versions differ, not the ones where both are obviously fine.
- Cap batch length; late-session labels are worse and nobody notices
- Rotate gold cases with known answers through batches to detect drift
- Blind and randomise presentation order — position effects are real
- Review where versions differ; identical-and-fine cases teach you nothing
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.