The AI Learning Hub Journal

Testing Around Non-Determinism

Testing something that answers differently every timeOne fixed inputsame prompt, same model,same tools, nothing changedrun it again and againThe spread of outputs you actually getTOLERANCE BAND — VARIATION YOU ACCEPToutlieroutlierterser, thinner, fewer stepslonger, more elaborate, more stepsTHE TEST THAT WILL NOT WORKAsserting one exact output string from one run: it fails on harmless rewording,and passes confident nonsense that happens to match.WHAT TO DO INSTEADRun it many timesOne run is an anecdote. Report apass rate across repeated runs andtreat a single green as noise.Assert invariantsValid schema, required fields, theright tool called, no leaked context— not the exact wording it chose.Use tolerance bandsAccept a range, and alert when thedistribution drifts rather than whenone run falls outside the band.Pin what can be pinnedFix temperature and seed, freeze themodel version and prompt, and logall of them beside the result.A flaky suite is usually a suite asserting on wording — pin what you can, and assert on what has to be true
The spread is the system behaving normally, not a bug — so the suite has to test the distribution, not one lucky sample

Where the Variance Comes From

Non-determinism in these systems has several independent sources, and knowing which one you are facing determines what to do about it. Sampling is the obvious one and the only one you control directly. Below it sits numerical non-determinism: floating-point reduction order varies with batching and kernel selection on GPUs, so even greedy decoding is not bitwise reproducible across a shared serving fleet, and providers generally do not guarantee otherwise. Then there is the environment — tools returning different data, timestamps, other users' concurrent activity, network flakiness. Finally, the provider itself: hosted models are updated, and behaviour can shift under a stable name. Temperature zero reduces variance substantially and does not eliminate it, which is why statistical treatment is the baseline rather than a fallback.

  • Sampling, floating-point batching effects, environment state, and provider-side change
  • Greedy decoding is not a determinism guarantee on shared infrastructure
  • Pin model versions where the platform allows and record them with every result
  • Design for variance rather than trying to engineer it away

Runs, Rates, and Where to Spend

Handle variance by running each scenario several times and asserting on the rate. Two framings are useful and answer different questions: the probability that a single attempt succeeds, and the probability that all of k attempts succeed — the latter matters when a user only sees one attempt and any failure is visible. For unattended automation, consistency across runs is often more important than peak capability, and reporting the worst observed run alongside the mean captures that. Given a fixed budget, the usual trade favours more cases over more runs per case, since case diversity samples the input distribution while extra runs only refine the estimate for inputs you already have — with the exception of scenarios you already know are borderline, where extra runs are exactly what resolves them.

  • Report pass rate, not a single transcript; a lucky run is not evidence
  • Consistency across runs matters more than peak for unattended automation
  • Budget usually buys more from additional cases than additional runs
  • Spend extra runs on scenarios already known to be marginal

Flaky Case or Degraded System

When a case starts failing intermittently you have to distinguish three explanations, and they call for opposite responses. The system genuinely regressed and the pass rate dropped — investigate and fix. The case was always marginal, sitting near the system's capability boundary, and normal variance is now visible — record its true rate and treat movement as signal. Or the case or environment is broken: ambiguous expectations, state leaking between runs, an unreliable external dependency. Historical pass rates settle this immediately, which is the argument for storing every run rather than only the latest verdict. A case whose rate has been stable at seventy percent for months and is stable at seventy percent today has not regressed, however uncomfortable the red mark in the report looks.

  • Store every run so historical rates are available at triage time
  • Marginal cases live near the capability boundary — they are informative, not broken
  • Order-dependent failures point at environment state, not the model
  • Quarantine genuinely broken cases; never delete a case that merely fails often

Try It Yourself

One run is a story, not a measurement. This is the cheapest way to find out how much your system moves on its own before you credit a change with moving it.

◆ Try it yourself

Take one scenario from your case sheet on a feature you ship, are building, or use daily — with no feature of your own, use a prompt you send often to an AI product you rely on. Write down what counts as a pass, then run the scenario ten times unchanged and count the passes. Now make one change you genuinely believe in — prompt wording, model, retrieval setting — and run the same scenario ten more times. Compare the two counts using the arithmetic below before you tell anyone the change worked. Reporting "too close to call" is a real result and is what keeps the number believable next time.

Scenario: [one real input]
Counts as a pass when: [written before any run]

Before the change: passes out of 10 = __
After the change:  passes out of 10 = __
Difference:        __ out of 10

How much wobble to expect: roughly 1 divided by the square root of the number of runs.
Ten runs: 1 divided by 3.16 is about 0.3 — so about 3 in 10.
A difference smaller than that is noise, not evidence.

Verdict: real move / too close to call / worth more runs
How you'll know it worked
  • The pass condition was written down before the first of the twenty runs
  • You have two counts out of ten and the difference written as a number, not two impressions
  • Your verdict is one of the honest options: a real move, or too close to call at ten runs

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.