The AI Learning Hub Journal
◆ Data

Sourcing Real Cases

Where eval cases come from, best source firstranked by how much a passing score on them tells you about live behaviour1Real production trafficActual inputs your users sent, drawn across the whole distributionthe only source that matches reality2Real failures and incidentsEverything that already went wrong, kept forever as a regression caseeach one is only paid for once3Support tickets and complaintsWhere users told you it was wrong, in their own words and contextskewed to the vocal, but still real4Expert-authored edge casesThe rare and awkward cases a domain expert already knows to fearcovers what traffic has not hit yet5Synthetic and generated casesModel-written variations that fill a coverage gap you can namegap filler, never the baselineTHE STEP THAT KEEPS THE SET REPRESENTATIVE RATHER THAN CHERRY-PICKED1Pool the candidatesEverything eligible fromthe sources above, beforeanybody picks favourites.2Stratify itSplit by what actuallychanges difficulty: type,length, language, user.3Sample within strataDraw from every stratum sorare-but-real slices arenot rounded away to zero.4Freeze and versionLock the set, record howit was drawn, and changeit only on purpose.Keep only the interesting cases and the score describes your taste, not the systemThe set has to look like the traffic, including the boring majority nobody wants to write downEvery incident is a free case — the expensive part already happened, so at least keep it
Cases you invent test the system you imagined — cases you collect test the one your users actually met

Production Is the Only Honest Source

Cases invented in a meeting room encode your assumptions about usage; cases pulled from logs encode usage. The difference shows up immediately — real inputs are messier, shorter, more ambiguous, full of typos, pasted formatting, missing context, and questions the product was never designed to answer. That mess is the distribution you are actually serving, and a suite built without it will report high scores on a system users find unreliable. So the first infrastructure decision in any AI product is logging designed for replay: the full input, the retrieved context, the assembled prompt, tool calls and results, the final output, and enough metadata to reconstruct the run. Without replayable logs you can collect complaints but you cannot build cases from them.

  • Log for replay, not just for debugging — capture context and assembled prompts
  • Real inputs are messier than imagined ones in ways that change results
  • Record the metadata you will want to slice on later: channel, segment, language, size

Sampling Deliberately

You cannot label everything, so how you sample determines what your suite can see. Random sampling gives you an unbiased picture of the common path and almost no coverage of rare failures. Stratified sampling across the slices you care about guarantees each one is represented well enough to measure. Failure-weighted sampling targets sessions with implicit distress signals — retries, rephrasings, abandonment, escalation to a human, thumbs-down — which is where the highest-value cases concentrate. Novelty sampling picks inputs unlike anything already in the set, using embedding distance, and is how you stop the suite calcifying around last quarter's traffic. Run all four with a fixed budget for each, and record which strategy produced each case so you can reason about the resulting mix rather than inheriting it blindly.

  • Random for the baseline, stratified for coverage, failure-weighted for value
  • Implicit signals — retry, rephrase, abandon, escalate — beat explicit feedback in volume
  • Novelty sampling by embedding distance keeps the set from calcifying
  • Tag each case with its sampling origin; a set of unknown composition cannot be interpreted

Synthetic Cases and Their Ceiling

Synthetic generation earns its place in three specific situations: before launch when no traffic exists, for rare-but-critical scenarios that real data will not supply in useful numbers, and for systematic perturbation of existing cases — the same request in another language, with a typo, with an unusual format, at ten times the length. What synthetic data does not do is tell you what users will actually ask, because it is generated from the same assumptions your prompt already encodes and inherits the generating model's stylistic distribution. A suite that is mostly synthetic tends to be clean, uniform, and quietly easier than reality. Use synthetic cases to fill known gaps, mark them clearly, and treat replacing them with real cases as ongoing work rather than an optional cleanup.

  • Good for cold start, rare critical scenarios, and systematic perturbation
  • Bad at anticipating real user intent — it inherits your assumptions
  • Label synthetic cases explicitly and track what proportion of the set they are
  • Perturbing real cases is usually higher value than generating fresh ones

Privacy, Consent, and Handling

Production data is the best eval source and the one with legal weight attached. Before a single real transcript lands in a repository, settle four questions: does your privacy notice and lawful basis cover secondary use for testing, what is redacted or pseudonymised on the way in, who can read the eval set, and how long cases are retained. Detect-and-redact pipelines for personal data are standard and imperfect, so pair them with access controls rather than trusting them alone. Where the domain is regulated, expect that constructing minimally realistic surrogates — same structure and difficulty, synthetic identifiers — is the only viable route, and budget the extra work. The failure to avoid is the common one: a golden dataset of genuine customer records sitting in a public source repository because nobody asked the question early.

  • Confirm lawful basis and notice coverage for secondary use before collecting
  • Redact or pseudonymise on ingest, then still control access to the set
  • Set retention rules for cases the same way you would for the source logs
  • In regulated domains, plan for structure-preserving surrogates from the start

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.