The AI Learning Hub Journal

Agreement, Significance, and Sample Size

Why an eval number needs an interval around itthe same comparison, measured twice — once on a small sample, once on a larger oneSMALL n — WIDE INTERVALthe two intervals overlap heavilyMODEL AMODEL Blower scorehigher scoreYou cannot tell these two apart yetLARGER n — INTERVAL NARROWSsame two systems, more samplesMODEL AMODEL Blower scorehigher scoreNow the gap clears both intervalsA difference smaller than the interval is not a result — it is the sample size talkingCHANCE-CORRECTED AGREEMENTTwo raters will agree some of the time by luck alone.Raw agreement quietly counts that luck as success.A chance-corrected figure subtracts what randommarking would have handed you anyway.PAIRED COMPARISON FOR A/B CHANGESRun both versions over the same items, rather thantwo separately drawn samples.Compare the per-item differences, so the shareddifficulty of the items cancels itself out.Report the interval and the sample size beside every score, or the score cannot be argued with
An eval score without an interval is a rumour — the interval is what tells you whether to act on it

Measuring Agreement Properly

Before trusting labels — human or model — measure how consistently they can be produced. Raw percentage agreement is the intuitive metric and it flatters badly on skewed data: if ninety-five percent of cases are passes, two raters who both label everything pass agree ninety-five percent of the time while conveying no information. Chance-corrected measures exist for exactly this. Cohen's kappa handles two raters on categorical labels, Fleiss' kappa extends to more, and Krippendorff's alpha copes with missing data and ordinal scales. Whichever you use, the number to care about is not an absolute threshold from a textbook but whether agreement is high enough that the differences you want to detect are larger than the noise in the labels.

  • Raw agreement is misleading whenever the label distribution is skewed
  • Cohen's kappa for two raters, Fleiss' for more, Krippendorff's alpha for messier designs
  • Low agreement means the rubric is ambiguous — fix wording, do not average the noise
  • Label noise sets a hard floor on the smallest effect your suite can resolve

Sample Size Is Arithmetic, Not Opinion

Teams routinely draw conclusions from twenty cases that the data cannot support. The arithmetic is straightforward: a measured pass rate around eighty percent on thirty cases carries a confidence interval of roughly plus or minus fourteen points, so a move from eighty to eighty-five is entirely consistent with nothing having changed. Precision improves with the square root of sample size, which means halving the interval requires quadrupling the cases — the reason detecting small deltas demands hundreds. Two practical consequences follow. Match your sample size to the effect size you need to detect, deciding what difference would actually change a decision before you run anything. And when a result is too close to call, say so; reporting a two-point gain from a small set as an improvement is how eval suites lose credibility.

  • Compute a confidence interval for every rate you report — a bare percentage overstates certainty
  • Precision scales with the square root of n: four times the cases for half the interval
  • Decide the effect size worth detecting before choosing the sample size
  • "Too close to call" is a legitimate, valuable result

Compare Variants the Cheap Way

The most effective way to get more statistical power without more cases is to compare variants on the same cases rather than on separate samples. Paired comparison removes case difficulty from the equation: you are no longer asking whether one sample scored higher, but on how many individual cases the outcome flipped and in which direction. Because most cases either pass or fail under both variants, the informative signal concentrates in the disagreements, and paired analysis reads them directly. Bootstrap resampling over the case set gives an interval for the difference without distributional assumptions. And watch multiple comparisons: evaluating ten prompt variants and celebrating the best one is a procedure that produces a winner from pure noise reliably enough that you should assume it has, unless the margin is large or the result replicates on held-out cases.

  • Always compare variants on identical cases — pairing is free statistical power
  • The signal lives in cases where the two variants disagree
  • Bootstrap the difference to get an interval without distributional assumptions
  • Testing many variants manufactures winners — confirm the leader on held-out cases

When Human Review Is Unavoidable

Automation has limits that no judge design removes, and recognising them is a mark of maturity rather than a failure of engineering. Human review is required when the stakes make an undetected error unacceptable — clinical, legal, financial, or safety-critical outputs. It is required for genuinely novel domains where no ground truth exists yet, because a judge can only apply a rubric someone has already written. It is required for the subjective heart of a product — whether a response is actually useful to this person in this situation — and for regulatory contexts where a human decision-maker is the point. The workable pattern is layered: automated grading over everything for coverage and trend, human review over a stratified sample for calibration, and mandatory human review on the highest-consequence slice, with disagreements between the two feeding straight back into the rubric.

  • High stakes, novel domains, deep subjectivity, and regulatory review resist automation
  • Layer it: automate for coverage, sample for calibration, mandate review where it counts
  • Route every human-judge disagreement back into the rubric — that is the improvement loop
  • Budget reviewer time as a permanent operating cost, not a one-off project expense

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.