Agreement, Significance, and Sample Size
Measuring Agreement Properly
Before trusting labels — human or model — measure how consistently they can be produced. Raw percentage agreement is the intuitive metric and it flatters badly on skewed data: if ninety-five percent of cases are passes, two raters who both label everything pass agree ninety-five percent of the time while conveying no information. Chance-corrected measures exist for exactly this. Cohen's kappa handles two raters on categorical labels, Fleiss' kappa extends to more, and Krippendorff's alpha copes with missing data and ordinal scales. Whichever you use, the number to care about is not an absolute threshold from a textbook but whether agreement is high enough that the differences you want to detect are larger than the noise in the labels.
- Raw agreement is misleading whenever the label distribution is skewed
- Cohen's kappa for two raters, Fleiss' for more, Krippendorff's alpha for messier designs
- Low agreement means the rubric is ambiguous — fix wording, do not average the noise
- Label noise sets a hard floor on the smallest effect your suite can resolve
Sample Size Is Arithmetic, Not Opinion
Teams routinely draw conclusions from twenty cases that the data cannot support. The arithmetic is straightforward: a measured pass rate around eighty percent on thirty cases carries a confidence interval of roughly plus or minus fourteen points, so a move from eighty to eighty-five is entirely consistent with nothing having changed. Precision improves with the square root of sample size, which means halving the interval requires quadrupling the cases — the reason detecting small deltas demands hundreds. Two practical consequences follow. Match your sample size to the effect size you need to detect, deciding what difference would actually change a decision before you run anything. And when a result is too close to call, say so; reporting a two-point gain from a small set as an improvement is how eval suites lose credibility.
- Compute a confidence interval for every rate you report — a bare percentage overstates certainty
- Precision scales with the square root of n: four times the cases for half the interval
- Decide the effect size worth detecting before choosing the sample size
- "Too close to call" is a legitimate, valuable result
Compare Variants the Cheap Way
The most effective way to get more statistical power without more cases is to compare variants on the same cases rather than on separate samples. Paired comparison removes case difficulty from the equation: you are no longer asking whether one sample scored higher, but on how many individual cases the outcome flipped and in which direction. Because most cases either pass or fail under both variants, the informative signal concentrates in the disagreements, and paired analysis reads them directly. Bootstrap resampling over the case set gives an interval for the difference without distributional assumptions. And watch multiple comparisons: evaluating ten prompt variants and celebrating the best one is a procedure that produces a winner from pure noise reliably enough that you should assume it has, unless the margin is large or the result replicates on held-out cases.
- Always compare variants on identical cases — pairing is free statistical power
- The signal lives in cases where the two variants disagree
- Bootstrap the difference to get an interval without distributional assumptions
- Testing many variants manufactures winners — confirm the leader on held-out cases
When Human Review Is Unavoidable
Automation has limits that no judge design removes, and recognising them is a mark of maturity rather than a failure of engineering. Human review is required when the stakes make an undetected error unacceptable — clinical, legal, financial, or safety-critical outputs. It is required for genuinely novel domains where no ground truth exists yet, because a judge can only apply a rubric someone has already written. It is required for the subjective heart of a product — whether a response is actually useful to this person in this situation — and for regulatory contexts where a human decision-maker is the point. The workable pattern is layered: automated grading over everything for coverage and trend, human review over a stratified sample for calibration, and mandatory human review on the highest-consequence slice, with disagreements between the two feeding straight back into the rubric.
- High stakes, novel domains, deep subjectivity, and regulatory review resist automation
- Layer it: automate for coverage, sample for calibration, mandate review where it counts
- Route every human-judge disagreement back into the rubric — that is the improvement loop
- Budget reviewer time as a permanent operating cost, not a one-off project expense
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.