The AI Learning Hub Journal

Golden Datasets and Curation

Curating the golden eval setPRODUCTION TRAFFICEverything real users sendLong tail, repeats, typos, off-topicSAMPLE, DO NOT SCRAPEStratify by intent and difficultyMirror the real usage mixLABEL EACH CASEInput · expected behaviour · grading ruleTwo labellers; resolve disagreements, do not averagesplit once, by case — never leak near-duplicates across the lineDEV SET · ~70%You look at it constantlyTune prompts, retrieval and models hereExpect to overfit it — that is its jobHELD-OUT SET · ~30%Locked. Never tuned against.Run it before a release, and rarelyOptimise against it once and it is a dev setCOVERAGE = THE REAL DISTRIBUTION + THE EDGES YOU ALREADY KNOW ABOUTReal mixmatch traffic proportionsKnown edgesadversarial, empty, huge, multilingualCostly casesover-sample what hurts mostFresh casesadd as the product and users shiftA set that does not look like production measures a product you are not shipping
A golden set is curated, not collected — and the held-out split is only worth anything for as long as nobody tunes against it

Anatomy of a Case

A case is more than an input and an expected answer. To be replayable and interpretable it needs the input as the user supplied it, any environment or state the system depends on, the expected outcome expressed in whatever form is gradable — a reference answer, a set of required facts, a rubric, a post-condition on the world — plus the grader to apply, slice tags, a difficulty marker, and provenance recording where the case came from and who labelled it. Provenance is the field teams skip and later need most: when a case is disputed, knowing whether it came from a production incident, a domain expert, or a generation script determines how much authority it carries. Store cases as versioned structured data in the repository, not in a spreadsheet whose edit history is a mystery.

  • Input, environment, expected outcome, grader, tags, difficulty, provenance
  • Expected outcome can be a reference, a fact list, a rubric, or a post-condition
  • Provenance settles disputes about whether a failing case is the system's fault
  • Version-controlled structured files beat spreadsheets for anything you will regenerate

Curation Is Ongoing Work

Sets rot. Cases go stale when the product changes, when the underlying data they reference is updated, or when the correct answer moves. Duplicates accumulate as similar production failures get filed repeatedly, silently over-weighting whatever failure mode was fashionable last month. Saturated cases — passed by every candidate system — consume compute and contribute nothing. And some cases are simply wrong: the expected answer was mislabelled, and until someone rechecks it, every model that answers correctly is scored as failing. Schedule curation rather than hoping for it. Deduplicate by embedding similarity, re-verify a sample of expected outputs each cycle, retire saturated cases, and be particularly suspicious of any case that every model fails — that is far more often a broken case than a universally hard task.

  • Deduplicate by similarity; repeated filings quietly reweight your metric
  • Retire saturated cases and promote newly discriminative ones
  • A case every system fails is usually mislabelled — check before treating it as a target
  • Re-verify a rotating sample of expected outputs; ground truth decays

Development and Held-Out Splits

The moment you start iterating against a set, you begin fitting to it — reading its failures, tuning prompts around its quirks, and improving on it faster than on reality. The standard defence transfers directly from machine learning: split. A development set is what you iterate against daily, inspect freely, and debug with. A held-out set is run rarely, at decision points, by someone who is not tuning against it, and is never read case by case for prompt inspiration. When development and held-out scores diverge, the gap is a direct measure of how much of your recent progress was overfitting. Refresh both from production periodically, and if the held-out set has been examined in detail, accept that it has been burned and rotate in a fresh one.

  • Iterate on dev, decide on held-out, and keep the roles strictly separate
  • A widening dev-versus-held-out gap is your overfitting signal
  • Reading the held-out set in detail spends it — plan for rotation
  • Refresh both splits from production so they track the current distribution

How Big, and How to Grow

The right size is driven by what you need to detect, not by a round number. Small sets — tens of cases — support fast iteration and catching gross breakage, and cannot resolve small differences. Detecting a few points of change reliably requires hundreds, and the arithmetic in the next lesson explains why. The practical path is to start with a few dozen carefully labelled real cases per critical slice, use them immediately, and grow from production evidence: every incident becomes a case, every recurring complaint becomes a case, every new feature ships with cases. Quality dominates quantity at every size, because a thousand cases with sloppy expected outputs produce a confident number that means nothing at all.

  • Start at a few dozen real cases per critical slice and ship them into CI immediately
  • Small sets catch breakage; resolving small deltas needs hundreds
  • Grow from incidents and complaints, not from generation scripts
  • Sloppy labels at scale yield confident, meaningless numbers

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.