Sourcing Real Cases
Production Is the Only Honest Source
Cases invented in a meeting room encode your assumptions about usage; cases pulled from logs encode usage. The difference shows up immediately — real inputs are messier, shorter, more ambiguous, full of typos, pasted formatting, missing context, and questions the product was never designed to answer. That mess is the distribution you are actually serving, and a suite built without it will report high scores on a system users find unreliable. So the first infrastructure decision in any AI product is logging designed for replay: the full input, the retrieved context, the assembled prompt, tool calls and results, the final output, and enough metadata to reconstruct the run. Without replayable logs you can collect complaints but you cannot build cases from them.
- Log for replay, not just for debugging — capture context and assembled prompts
- Real inputs are messier than imagined ones in ways that change results
- Record the metadata you will want to slice on later: channel, segment, language, size
Sampling Deliberately
You cannot label everything, so how you sample determines what your suite can see. Random sampling gives you an unbiased picture of the common path and almost no coverage of rare failures. Stratified sampling across the slices you care about guarantees each one is represented well enough to measure. Failure-weighted sampling targets sessions with implicit distress signals — retries, rephrasings, abandonment, escalation to a human, thumbs-down — which is where the highest-value cases concentrate. Novelty sampling picks inputs unlike anything already in the set, using embedding distance, and is how you stop the suite calcifying around last quarter's traffic. Run all four with a fixed budget for each, and record which strategy produced each case so you can reason about the resulting mix rather than inheriting it blindly.
- Random for the baseline, stratified for coverage, failure-weighted for value
- Implicit signals — retry, rephrase, abandon, escalate — beat explicit feedback in volume
- Novelty sampling by embedding distance keeps the set from calcifying
- Tag each case with its sampling origin; a set of unknown composition cannot be interpreted
Synthetic Cases and Their Ceiling
Synthetic generation earns its place in three specific situations: before launch when no traffic exists, for rare-but-critical scenarios that real data will not supply in useful numbers, and for systematic perturbation of existing cases — the same request in another language, with a typo, with an unusual format, at ten times the length. What synthetic data does not do is tell you what users will actually ask, because it is generated from the same assumptions your prompt already encodes and inherits the generating model's stylistic distribution. A suite that is mostly synthetic tends to be clean, uniform, and quietly easier than reality. Use synthetic cases to fill known gaps, mark them clearly, and treat replacing them with real cases as ongoing work rather than an optional cleanup.
- Good for cold start, rare critical scenarios, and systematic perturbation
- Bad at anticipating real user intent — it inherits your assumptions
- Label synthetic cases explicitly and track what proportion of the set they are
- Perturbing real cases is usually higher value than generating fresh ones
Privacy, Consent, and Handling
Production data is the best eval source and the one with legal weight attached. Before a single real transcript lands in a repository, settle four questions: does your privacy notice and lawful basis cover secondary use for testing, what is redacted or pseudonymised on the way in, who can read the eval set, and how long cases are retained. Detect-and-redact pipelines for personal data are standard and imperfect, so pair them with access controls rather than trusting them alone. Where the domain is regulated, expect that constructing minimally realistic surrogates — same structure and difficulty, synthetic identifiers — is the only viable route, and budget the extra work. The failure to avoid is the common one: a golden dataset of genuine customer records sitting in a public source repository because nobody asked the question early.
- Confirm lawful basis and notice coverage for secondary use before collecting
- Redact or pseudonymise on ingest, then still control access to the set
- Set retention rules for cases the same way you would for the source logs
- In regulated domains, plan for structure-preserving surrogates from the start
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.