The AI Learning Hub Journal

Human Review That Scales

Human review never reaches zero — spend it where it discriminatesautomate the checkable, delegate the articulable to judges — what remains is judgement, and it is the scarce inputEVERY OUTPUT, THREE TIERSautomated gradersanything with a checkable definitionmodel judgescriteria you can write downhuman tierirreducible judgementa permanent tier — it shrinks asrubrics improve, never to zeroappropriateness, tone, safety incontext, whether it actually helpsSAMPLE DELIBERATELY, NOT UNIFORMLYa small random sample keeps the base rate honeststratify on top: high-stakes slices, borderline cases,grader disagreements, and verdict flips between versionsuniform sampling re-confirms what automation already knewrecord the sampling scheme beside every number it producesTWO REVIEWERS DISAGREE MORE THAN THEY EXPECTdouble-label part of every batch; track agreement as a metriclow agreement is a rubric defect — swap adjectives forobservable conditions, anchor scale points with exampleshuman agreement caps any judge validated against those labelsFATIGUE AND DRIFT — COUNTER THEM MECHANICALLYcap batch length —late labels are worserotate gold cases throughbatches to catch driftblind and randomise theorder — position effectsre-run an old batch —do labels still hold?REVIEW WHERE VERSIONS DIFFER — IDENTICAL-AND-FINE CASES TEACH YOU NOTHINGA PERMANENT TIER — SAMPLED, CALIBRATED, AND PROTECTED FROM FATIGUEthe process exists to allocate expert attention, not to review everything
Human review scales by sampling deliberately, measuring reviewer agreement, and countering fatigue — not by trying to review everything.

Where Judgement Cannot Be Delegated

Automated graders handle anything with a checkable definition of correct — schema validity, a computed figure, a required fact, a forbidden action — and model judges extend that to criteria you can write down explicitly. What remains is genuinely irreducible: whether an answer is appropriate for this customer in this situation, whether the tone fits, whether advice is safe given domain norms a rubric cannot fully capture, and whether the thing is actually useful rather than merely responsive. These are the cases where the rubric you would have to write is the judgement itself. Human review is therefore not a stage that teams graduate out of once they automate properly; it is a permanent tier that shrinks as your rubrics improve and never reaches zero. Treat expert attention as the scarcest input in the whole eval programme and design the process around spending it well.

  • Automate anything checkable; use judges for anything you can articulate
  • What is left is judgement a rubric cannot express — appropriateness, tone, real usefulness
  • Human review is a permanent tier, not a temporary one
  • Expert attention is the scarce resource; the process exists to allocate it

Sample, Do Not Review Everything

Reviewing every case is neither possible nor useful, and the instinct to try produces a process that gets quietly abandoned. Sample instead, and sample deliberately rather than uniformly. A small random sample gives an unbiased estimate and keeps you honest about the base rate. On top of that, stratify: over-sample the slices that matter most, the cases automated graders flagged as borderline, the cases where two graders disagreed, and the cases that changed verdict between the last version and this one. Uniform random sampling spends most of its budget confirming that easy cases still pass, which is precisely what automation already told you. Choose the sample size from the precision you need rather than from what feels thorough, and record which sampling scheme produced a number, because comparing a stratified review against last quarter's random one compares nothing at all.

  • A small random sample keeps the base rate honest; stratify on top of it
  • Over-sample high-stakes slices, borderline cases, and verdict changes
  • Uniform sampling re-confirms what automation already established
  • Record the sampling scheme with the number or cross-run comparisons are meaningless

Making Two Reviewers Agree

Two competent reviewers given the same output and the same one-line criterion will disagree far more than either expects, and unmeasured disagreement means your human numbers are noise with a serious face on. Measure it: have a portion of every batch labelled independently by two people and track the agreement rate as a standing metric. When it is low, the rubric is the problem rather than the reviewers — the fix is to convert vague adjectives into observable conditions, anchor each point of the scale with worked examples including the boundaries, and write explicit rules for the cases people keep resolving differently. Run a calibration session on a shared set before a batch, discuss the disagreements instead of averaging them away, and reissue the rubric. Agreement between humans is also the ceiling for any model judge validated against those labels, which makes this a prerequisite rather than a refinement.

  • Double-label part of every batch and track agreement as a standing metric
  • Low agreement is a rubric defect: replace adjectives with observable conditions
  • Anchor every scale point with worked examples, especially at the boundaries
  • Human agreement caps the accuracy of any judge trained against those labels

Fatigue, Drift, and Attention That Discriminates

Reviewer quality degrades in ways that show up as nothing on a dashboard. Long batches produce fatigue, and labels late in a session are measurably worse than labels early in it. Standards drift over weeks as reviewers recalibrate against whatever they have been seeing, so a score from one quarter and a score from the next are not directly comparable. Presentation order matters too, since whichever output appears first carries an advantage. Counter all three mechanically: cap batch length, rotate a small set of gold cases with known answers through every batch to detect drift, blind and randomise the order of candidates, and periodically re-run an old batch to check that today's reviewers still agree with the earlier labels. Then spend what remains where it discriminates — on the cases where two versions differ, not the ones where both are obviously fine.

  • Cap batch length; late-session labels are worse and nobody notices
  • Rotate gold cases with known answers through batches to detect drift
  • Blind and randomise presentation order — position effects are real
  • Review where versions differ; identical-and-fine cases teach you nothing

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.