Rubrics Humans and Models Can Both Apply
The Rubric Is the Specification
Writing a rubric feels like documentation and is actually specification work — it is where the vague ambition of a good answer becomes a set of testable statements. That is why rubric-writing surfaces disagreements the team did not know it had: two people who both want accurate summaries turn out to mean different things by accurate, and the rubric is where that gets resolved instead of relitigated in every review. Because the rubric doubles as the instruction set for both human reviewers and model judges, it has to be written for someone with no additional context: no references to internal conventions, no criteria that depend on knowing what the team discussed last month, and no words like appropriate or reasonable standing alone without a definition of what they mean here.
- Rubric-writing is where implicit quality standards become explicit and arguable
- Write for a competent stranger — no unexplained internal context
- Undefined words like appropriate, clear, and reasonable are where consistency dies
- The same rubric drives human labelling and model judging, so it must serve both
Anatomy of a Criterion
A criterion that survives contact with multiple raters has four properties. It is atomic, testing exactly one thing — compound criteria joined by and produce a single verdict on two questions and make disagreement impossible to diagnose. It is observable from the output plus its inputs, requiring no knowledge the rater does not have. It is decidable, phrased so a careful reader reaches a yes or no without weighing competing considerations. And it is anchored with examples: at minimum one clear pass, one clear fail, and one borderline case with the ruling and the reason. Anchor examples do more work than any amount of additional prose, because they convert an abstract standard into a demonstrated boundary — and they can be dropped straight into a judge prompt as few-shot guidance.
- Atomic: split every criterion containing "and" into separate checks
- Observable: graders can only see the output, the input, and the provided context
- Anchored: pass, fail, and borderline examples with the reasoning attached
- Anchor examples double as few-shot examples for a model judge
Iterate the Rubric on Disagreement
The first version of a rubric is always underspecified, and the way to find out where is to have two people apply it independently to the same twenty cases and then examine every disagreement. Each one is a defect in the rubric, not a defect in a rater: it marks a place where the wording admits two readings. Fix the wording, add an anchor covering the disputed situation, and rerun until agreement is high enough to be useful. Only then is the rubric ready for a model judge, because a rubric humans cannot apply consistently will produce noise from a model too — with the added hazard that the model produces its noise fluently and with a confident explanation attached. Rubric development is cheap at this stage and expensive after you have labelled a thousand cases with an ambiguous version.
- Two independent raters on the same small sample, then examine every disagreement
- Each disagreement is a rubric defect — fix wording and add an anchor
- Reach human consistency before automating; a judge inherits every ambiguity
- Freeze and version the rubric once labelling starts; changing it invalidates prior labels
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.