Reading Public Benchmarks Honestly
What They Are Good For
Public benchmarks are not worthless — they are just weaker evidence than their presentation implies. Used correctly they do three useful jobs. They track capability trends over time, which is how you know whether a class of problem has become tractable at all. They provide a coarse shortlist: a model far behind on reasoning-heavy suites is unlikely to surprise you on your reasoning-heavy task. And they give the field a shared vocabulary for discussing capability, without which every conversation restarts from scratch. What they cannot do is predict performance on your workload, because your workload has its own distribution, its own definition of correct, and a prompt and tool stack that no benchmark harness includes. Shortlist with public scores; decide with private ones.
- Good for capability trends, coarse shortlisting, and shared vocabulary
- Bad at predicting performance on any specific production workload
- Small leaderboard gaps near the top rarely transfer to your tasks
- Shortlist publicly, decide privately — never the reverse
Contamination
Benchmarks are published, models train on published text, so a high score may reflect memorisation rather than capability. This is structural rather than accidental, and it is hard to rule out: contamination can arrive through direct inclusion, through the countless derivative discussions, solutions, and tutorials that surround any popular benchmark, or through synthetic data generated by a model that had itself seen the test. The observable symptom is a gap between performance on a benchmark and performance on freshly constructed problems of the same type and difficulty. Defences exist and each is partial — canary strings that dataset authors ask trainers to filter, private held-out splits scored by the maintainer, and continuously refreshed benchmarks built from material published after training cutoffs. Treat any single reported score as an upper bound.
- Contamination arrives indirectly too — via solutions, discussion, and synthetic derivatives
- The tell is a gap between benchmark scores and fresh problems of matched difficulty
- Private held-out splits and continuously refreshed sets are partial defences
- Read reported scores as upper bounds, especially on older, popular suites
Saturation, Construct Validity, and Arenas
Two further problems limit what a benchmark number can mean. Saturation: once the frontier clusters near the ceiling, remaining headroom is dominated by label noise and ambiguous items, so ranking differences at the top measure error more than capability. Construct validity: a benchmark measures a specific operationalisation of a skill, and multiple-choice knowledge questions or self-contained coding problems are structurally unlike open-ended work in a real environment with ambiguous requirements and messy context. Human-preference arenas fix some of this by using real prompts and real comparisons, but they measure what raters prefer — which rewards style, length, and formatting alongside correctness — and their voter population is not your user population. Every benchmark is a proxy; know which proxy you are reading.
- Near the ceiling, ranking deltas mostly reflect label noise
- Exam-shaped tasks are not product-shaped tasks — beware construct validity
- Preference arenas measure preference, including style and length effects
- Agentic benchmarks depend heavily on scaffold and environment, so scores travel poorly
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.