Testing Around Non-Determinism
Where the Variance Comes From
Non-determinism in these systems has several independent sources, and knowing which one you are facing determines what to do about it. Sampling is the obvious one and the only one you control directly. Below it sits numerical non-determinism: floating-point reduction order varies with batching and kernel selection on GPUs, so even greedy decoding is not bitwise reproducible across a shared serving fleet, and providers generally do not guarantee otherwise. Then there is the environment — tools returning different data, timestamps, other users' concurrent activity, network flakiness. Finally, the provider itself: hosted models are updated, and behaviour can shift under a stable name. Temperature zero reduces variance substantially and does not eliminate it, which is why statistical treatment is the baseline rather than a fallback.
- Sampling, floating-point batching effects, environment state, and provider-side change
- Greedy decoding is not a determinism guarantee on shared infrastructure
- Pin model versions where the platform allows and record them with every result
- Design for variance rather than trying to engineer it away
Runs, Rates, and Where to Spend
Handle variance by running each scenario several times and asserting on the rate. Two framings are useful and answer different questions: the probability that a single attempt succeeds, and the probability that all of k attempts succeed — the latter matters when a user only sees one attempt and any failure is visible. For unattended automation, consistency across runs is often more important than peak capability, and reporting the worst observed run alongside the mean captures that. Given a fixed budget, the usual trade favours more cases over more runs per case, since case diversity samples the input distribution while extra runs only refine the estimate for inputs you already have — with the exception of scenarios you already know are borderline, where extra runs are exactly what resolves them.
- Report pass rate, not a single transcript; a lucky run is not evidence
- Consistency across runs matters more than peak for unattended automation
- Budget usually buys more from additional cases than additional runs
- Spend extra runs on scenarios already known to be marginal
Flaky Case or Degraded System
When a case starts failing intermittently you have to distinguish three explanations, and they call for opposite responses. The system genuinely regressed and the pass rate dropped — investigate and fix. The case was always marginal, sitting near the system's capability boundary, and normal variance is now visible — record its true rate and treat movement as signal. Or the case or environment is broken: ambiguous expectations, state leaking between runs, an unreliable external dependency. Historical pass rates settle this immediately, which is the argument for storing every run rather than only the latest verdict. A case whose rate has been stable at seventy percent for months and is stable at seventy percent today has not regressed, however uncomfortable the red mark in the report looks.
- Store every run so historical rates are available at triage time
- Marginal cases live near the capability boundary — they are informative, not broken
- Order-dependent failures point at environment state, not the model
- Quarantine genuinely broken cases; never delete a case that merely fails often
Try It Yourself
One run is a story, not a measurement. This is the cheapest way to find out how much your system moves on its own before you credit a change with moving it.
Take one scenario from your case sheet on a feature you ship, are building, or use daily — with no feature of your own, use a prompt you send often to an AI product you rely on. Write down what counts as a pass, then run the scenario ten times unchanged and count the passes. Now make one change you genuinely believe in — prompt wording, model, retrieval setting — and run the same scenario ten more times. Compare the two counts using the arithmetic below before you tell anyone the change worked. Reporting "too close to call" is a real result and is what keeps the number believable next time.
Scenario: [one real input] Counts as a pass when: [written before any run] Before the change: passes out of 10 = __ After the change: passes out of 10 = __ Difference: __ out of 10 How much wobble to expect: roughly 1 divided by the square root of the number of runs. Ten runs: 1 divided by 3.16 is about 0.3 — so about 3 in 10. A difference smaller than that is noise, not evidence. Verdict: real move / too close to call / worth more runs
- The pass condition was written down before the first of the twenty runs
- You have two counts out of ten and the difference written as a number, not two impressions
- Your verdict is one of the honest options: a real move, or too close to call at ten runs
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.