The AI Learning Hub Journal

LLM-as-Judge, Done Properly

LLM-as-judge and its known failure modesResponse AResponse BLLM-AS-JUDGEscores an answer, or picks a winnercheap, fast, and it scales to every runbut it is a model, with model failure modesVerdictA wins · 4 / 5POSITION BIASPrefers whichever answer itreads first — or last.MITIGATIONSwap the order, run twice, average.VERBOSITY BIASRewards length and confidenttone over being correct.MITIGATIONScore against an explicit rubric.SELF-PREFERENCERates text from its own modelfamily above everyone else.MITIGATIONJudge with a different model family.CALIBRATION — the step that turns a judge into a measurement1 · HUMAN-LABEL A SLICE~100 real cases,labelled by people2 · RUN THE JUDGEsame slice,same rubric3 · MEASURE AGREEMENTCohen kappa, notraw accuracy4 · TIGHTEN AND REPEATrewrite the rubricuntil it clears your barRecalibrate whenever the rubric, the judge model or the task changes — agreement does not carry over
An LLM judge is a model under test, not a measuring instrument — until you have measured its agreement with humans

Why Judges, and What They Actually Are

Many properties worth measuring resist code — whether a summary is faithful to its source, whether an explanation would satisfy the person who asked, whether a refusal was warranted. Human review answers these and does not scale to every change on every pull request. A model judge closes the gap and now underpins most serious eval pipelines. The framing that keeps teams honest is to treat the judge as a measuring instrument: a device with a resolution, a bias profile, and a calibration state. You would not trust an uncalibrated instrument in any other engineering discipline, and an uncalibrated judge is worse than no metric because it produces a confident number with an articulate justification attached, which is far harder to disbelieve than silence.

  • Judges exist for properties no assertion can express — not as a default grader
  • Treat a judge as an instrument with bias and calibration, not an oracle
  • A fluent wrong grade is more dangerous than a missing grade
  • Every judge is a second AI system in your stack, with its own failure modes

Prompting a Judge Well

Judge quality is mostly prompt design, and the same handful of choices account for most of the gap between a useful judge and a random number generator. Give it the rubric with its anchor examples rather than a bare instruction. Ask for one criterion at a time; a judge asked for six dimensions in one call tends to produce correlated scores that mostly reflect a general impression. Require a short justification before the verdict, not after, so the label is conditioned on the reasoning rather than rationalising a snap decision. Demand structured output so parsing never becomes a second source of error. Provide reference answers where they exist, since grading against a reference is a far easier task than grading in the abstract. And avoid open numeric scales — a request to rate something out of ten produces clustering and drift that no amount of prompt tuning removes.

  • One criterion per call, with the rubric and anchors included
  • Reasoning first, then a structured verdict — order matters
  • Reference answers turn an open judgement into a comparison
  • Binary or few discrete levels; open 1–10 scales drift and cluster

The Bias Profile

Judge biases are documented, reproducible, and largely mitigable if you design for them. Position bias: in pairwise comparisons, judges systematically favour one position, so run both orderings and treat order-dependent verdicts as ties rather than data. Verbosity bias: longer answers score higher at equal quality, so control for length in comparisons and watch whether your outputs are getting longer as scores improve. Self-preference: models favour text resembling their own generation style, so avoid grading a family with itself where you can. Beyond these, judges reward confident phrasing over hedged accuracy, are influenced by formatting and markdown structure, and can be swayed by assertions inside the content being graded — which makes a judge one more component with a prompt-injection surface.

  • Position bias: always swap orderings; disagreement across orders means tie
  • Verbosity bias: control for length, and monitor output length as a side effect of tuning
  • Self-preference: judge with a different model family than you generate with where feasible
  • Judges read attacker-influenced content — treat graded text as untrusted input

Calibration Against Human Labels

A judge is only a metric once you have measured it against the ground truth it is standing in for. The procedure is unglamorous and non-negotiable: take a sample of cases, have your domain expert label them by the same rubric, and compare — not with a single accuracy figure but as a classification problem. Where does the judge produce false passes, where false failures, and is the error systematic on a particular slice? Compare its agreement with humans against how well two humans agree with each other, because that human ceiling is the realistic upper bound and a judge close to it is doing well. Then re-validate whenever anything changes: a new judge model, a rubric edit, a shift in output distribution. Provider-side model updates can move judge behaviour without any change on your side, which is why holding the judge version stable and re-checking periodically matters.

  • Compare judge to human labels as a confusion matrix, not one accuracy number
  • Benchmark judge-human agreement against human-human agreement, the realistic ceiling
  • Re-validate on every judge change, rubric edit, or major output shift
  • Pin the judge model where the platform allows; silent updates move your metric

Try It Yourself

Calibration sounds like a project and is really one afternoon with twenty cases. It is also the only step that turns a judge score from an opinion into a measurement.

◆ Try it yourself

Take twenty outputs from an AI feature you ship, are building, or use daily — with none of your own, collect twenty answers from an AI product you rely on. Choose one binary criterion from your rubric, label all twenty yourself first, then run the judge prompt below over the same twenty and compare. Count the disagreements in two directions: how many the judge passed and you failed, and how many it failed and you passed. Then read every disagreement and say what wording in the criterion allowed two honest readings — almost all of them are a rubric bug, not a judge bug.

You are grading one criterion only.

Criterion: [one yes/no criterion, in observable terms]
Clearly passes: [an example, and why]
Clearly fails: [an example, and why]
Borderline: [a case near the line, with the ruling and the reason]

Input the system was given:
[paste]

Output to grade:
[paste]

Write two sentences of reasoning first. Then put the verdict on the last line, as exactly PASS or FAIL.
How you'll know it worked
  • Each of the twenty items carries two recorded labels — yours, written before you saw the judge's
  • The disagreements are reported as two counts, not one: judge-passed-you-failed and judge-failed-you-passed
  • At least one disagreement produced an edit to the criterion wording or a new anchor example

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.