The AI Learning Hub Journal
◆ Operations

Non-Determinism as an Engineering Constraint

Two runs of the same task will legitimately differper-call sampling variance compounds into divergent trajectories — engineer for outcomes and invariants, not the pathONE TASK, THREE LEGITIMATE RUNS — JUDGED ONLY WHERE THEY ENDthe same task,the same prompt4 steps, done6 steps, done5 steps, doneOUTCOMES ANDINVARIANTSthe only stablethings to judgereducing randomness narrows the spread but does not remove it — and any prompt or index change reshuffles behaviourMAKE RE-RUNNING SAFEretry is the main recoverypath, so tool idempotencyis a reliability feature,not only a correctness oneENFORCE INVARIANTSnever writes outside theserecords, never exceeds thisspend — checked in theruntime, not trustedPREFER CHEAP RETRIESover expensive attemptsthat must thereforesucceed — failure becomesa handled, ordinary eventPIN WHAT YOU CANmodel, prompt, schema andindex versions recorded onevery run — to tell whatmoved when behaviour shiftsONE SUCCESSFUL TRY IS NOT EVIDENCEthe engineer runs the new prompt once, it works, it ships — and the regression ships with itJUDGE CHANGES ON DISTRIBUTIONS — PASS RATE AGAINST PASS RATEbuild repeatable execution in from the start: environment reset, seeded fixtures, a runner over stored cases
Non-determinism is survived, not eliminated — judge outcomes and invariants, and evaluate changes on distributions

The Same Input Will Not Give the Same Run

Sampling makes any single model call variable, and a multi-step loop compounds that variability into genuinely divergent trajectories: two runs of the same task can differ in which tools they call, how many steps they take and what they conclude. Reducing sampling randomness narrows this but does not remove it, because providers do not guarantee identical outputs across time or infrastructure, and any change to a prompt, a tool description or a retrieved document legitimately reshuffles behaviour. The engineering response is not to chase determinism. It is to stop depending on it: design so that correctness is a property of the outcome and the invariants rather than of the path, and so that any single run failing is an expected event the system handles rather than an anomaly that requires a person.

  • Per-call variance compounds into divergent trajectories over a multi-step run
  • Lowering randomness narrows the spread; it does not give you reproducibility
  • Correctness must be a property of outcomes and invariants, not of the path taken
  • A single failed run should be a handled event, not an exception requiring a human

Design So Variance Is Survivable

Several concrete practices follow. Make re-running safe, which means the idempotency work in the tool layer is a reliability requirement and not only a correctness one, because your recovery strategy for a large share of failures will be to try again. Define invariants that must hold on every run — this class of agent never writes outside these records, never exceeds this spend, always produces a result carrying these fields — and enforce them in the runtime where they are checkable rather than trusting them to hold. Prefer designs where a second attempt is cheap over designs where each attempt is expensive and must therefore succeed. And pin what you actually can: model family and version, prompt version, tool schema version, retrieval index version, all recorded on the run, so that when behaviour shifts you can tell whether anything under your control moved.

  • Idempotency is a reliability requirement, because retry is your main recovery path
  • Enforce per-run invariants in the runtime rather than trusting them to hold
  • Prefer cheap-retry designs over designs where each attempt must succeed
  • Record model, prompt, schema and index versions on every run

Judging Whether Something Changed

Because runs vary, no single run is evidence about a change, and this is where teams most often mislead themselves: an engineer tries the new prompt once, it works, and it ships. Judgements about behaviour require repetition and a distribution — a scenario run several times, with the pass rate compared against the pass rate before. That machinery is the subject of a neighbouring discipline and deserves its own treatment; the point here is architectural. Build the system so a scenario can be executed repeatedly without manual setup: environment reset, seeded fixtures, a runner that takes a stored case and produces a structured result. Teams that skip this end up unable to answer whether anything improved, and settle into changing prompts by intuition, which works until the second person joins the project.

  • One successful run is not evidence that a change helped
  • Behaviour claims need repeated runs and a pass-rate comparison
  • Build for repeatable execution: environment reset, fixtures, a case runner
  • Without it, prompt changes are made on intuition and regressions ship unseen

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.