The AI Learning Hub Journal

Rubrics Humans and Models Can Both Apply

A rubric a human and a model can both apply to the same outputnamed dimensions, concrete level descriptors, and a worked example anchoring each levelDIMENSIONS DOWN, LEVELS ACROSS — EVERY CELL SAYS WHAT THE TEXT DOES, NOT HOW GOOD IT FEELSLEVEL 1 — UNACCEPTABLELEVEL 2 — ACCEPTABLELEVEL 3 — EXEMPLARYFaithfulnessto the sourceContains a claim the sourcedoes not support.anchor: cites a figure not in the sourceEvery claim traceable to thesource; nothing invented.anchor: each point maps to a passageTraceable, and flags where thesource is silent.anchor: notes where the source is silentInstructioncomplianceIgnores a stated constrainton format, length or scope.anchor: asked for bullets, got proseMeets every statedconstraint.anchor: bullets, in the order askedMeets them, and asks when twoconstraints conflict.anchor: flags that two constraints clashCompletenessof coverageOmits a part of the requestwithout saying so.anchor: answers one of two questionsCovers every part of therequest.anchor: both questions answeredCovers it, and names what itdeliberately left out.anchor: names what it left outWHY A JUDGE MODEL AND A HUMAN DISAGREE — IT IS USUALLY THE WORDINGVAGUE CRITERION — “IS THE ANSWER GOOD?”Two raters read “good” differently and nothing in the rubricsettles it. The judge model inherits exactly the sameambiguity — so judge-human disagreement is a rubric defect.Same output, two raters, two different scores — and no way to settle it.CONCRETE DESCRIPTOR — WHAT THE TEXT DOESEach level names an observable property of the text, so tworaters check the same thing — and so does the judge. What isleft is real ambiguity in the output, not in the wording.Same output, two raters, the same score, checked the same way.THE TEST OF A RUBRIC IS AGREEMENT, NOT ELEGANCEScore it twice, independentlyTwo raters, no discussion,the same set of outputs.Read every disagreementEach one marks a place thedescriptor failed to decide.Fix the rubric, not the ratersRewrite the level until thetwo would have had to agree.A rubric works when two independent raters reach the same score without conferringVague criteria are exactly why a judge and a human disagree — fix the words, not the judgeWrite the levels so a stranger could score the same output and land in the same place
The rubric is the specification — if it does not decide the case, neither the human nor the model can

The Rubric Is the Specification

Writing a rubric feels like documentation and is actually specification work — it is where the vague ambition of a good answer becomes a set of testable statements. That is why rubric-writing surfaces disagreements the team did not know it had: two people who both want accurate summaries turn out to mean different things by accurate, and the rubric is where that gets resolved instead of relitigated in every review. Because the rubric doubles as the instruction set for both human reviewers and model judges, it has to be written for someone with no additional context: no references to internal conventions, no criteria that depend on knowing what the team discussed last month, and no words like appropriate or reasonable standing alone without a definition of what they mean here.

  • Rubric-writing is where implicit quality standards become explicit and arguable
  • Write for a competent stranger — no unexplained internal context
  • Undefined words like appropriate, clear, and reasonable are where consistency dies
  • The same rubric drives human labelling and model judging, so it must serve both

Anatomy of a Criterion

A criterion that survives contact with multiple raters has four properties. It is atomic, testing exactly one thing — compound criteria joined by and produce a single verdict on two questions and make disagreement impossible to diagnose. It is observable from the output plus its inputs, requiring no knowledge the rater does not have. It is decidable, phrased so a careful reader reaches a yes or no without weighing competing considerations. And it is anchored with examples: at minimum one clear pass, one clear fail, and one borderline case with the ruling and the reason. Anchor examples do more work than any amount of additional prose, because they convert an abstract standard into a demonstrated boundary — and they can be dropped straight into a judge prompt as few-shot guidance.

  • Atomic: split every criterion containing "and" into separate checks
  • Observable: graders can only see the output, the input, and the provided context
  • Anchored: pass, fail, and borderline examples with the reasoning attached
  • Anchor examples double as few-shot examples for a model judge

Iterate the Rubric on Disagreement

The first version of a rubric is always underspecified, and the way to find out where is to have two people apply it independently to the same twenty cases and then examine every disagreement. Each one is a defect in the rubric, not a defect in a rater: it marks a place where the wording admits two readings. Fix the wording, add an anchor covering the disputed situation, and rerun until agreement is high enough to be useful. Only then is the rubric ready for a model judge, because a rubric humans cannot apply consistently will produce noise from a model too — with the added hazard that the model produces its noise fluently and with a confident explanation attached. Rubric development is cheap at this stage and expensive after you have labelled a thousand cases with an ambiguous version.

  • Two independent raters on the same small sample, then examine every disagreement
  • Each disagreement is a rubric defect — fix wording and add an anchor
  • Reach human consistency before automating; a judge inherits every ambiguity
  • Freeze and version the rubric once labelling starts; changing it invalidates prior labels

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.