Non-Determinism as an Engineering Constraint
The Same Input Will Not Give the Same Run
Sampling makes any single model call variable, and a multi-step loop compounds that variability into genuinely divergent trajectories: two runs of the same task can differ in which tools they call, how many steps they take and what they conclude. Reducing sampling randomness narrows this but does not remove it, because providers do not guarantee identical outputs across time or infrastructure, and any change to a prompt, a tool description or a retrieved document legitimately reshuffles behaviour. The engineering response is not to chase determinism. It is to stop depending on it: design so that correctness is a property of the outcome and the invariants rather than of the path, and so that any single run failing is an expected event the system handles rather than an anomaly that requires a person.
- Per-call variance compounds into divergent trajectories over a multi-step run
- Lowering randomness narrows the spread; it does not give you reproducibility
- Correctness must be a property of outcomes and invariants, not of the path taken
- A single failed run should be a handled event, not an exception requiring a human
Design So Variance Is Survivable
Several concrete practices follow. Make re-running safe, which means the idempotency work in the tool layer is a reliability requirement and not only a correctness one, because your recovery strategy for a large share of failures will be to try again. Define invariants that must hold on every run — this class of agent never writes outside these records, never exceeds this spend, always produces a result carrying these fields — and enforce them in the runtime where they are checkable rather than trusting them to hold. Prefer designs where a second attempt is cheap over designs where each attempt is expensive and must therefore succeed. And pin what you actually can: model family and version, prompt version, tool schema version, retrieval index version, all recorded on the run, so that when behaviour shifts you can tell whether anything under your control moved.
- Idempotency is a reliability requirement, because retry is your main recovery path
- Enforce per-run invariants in the runtime rather than trusting them to hold
- Prefer cheap-retry designs over designs where each attempt must succeed
- Record model, prompt, schema and index versions on every run
Judging Whether Something Changed
Because runs vary, no single run is evidence about a change, and this is where teams most often mislead themselves: an engineer tries the new prompt once, it works, and it ships. Judgements about behaviour require repetition and a distribution — a scenario run several times, with the pass rate compared against the pass rate before. That machinery is the subject of a neighbouring discipline and deserves its own treatment; the point here is architectural. Build the system so a scenario can be executed repeatedly without manual setup: environment reset, seeded fixtures, a runner that takes a stored case and produces a structured result. Teams that skip this end up unable to answer whether anything improved, and settle into changing prompts by intuition, which works until the second person joins the project.
- One successful run is not evidence that a change helped
- Behaviour claims need repeated runs and a pass-rate comparison
- Build for repeatable execution: environment reset, fixtures, a case runner
- Without it, prompt changes are made on intuition and regressions ship unseen
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.