Production Reliability & Observability
Tracing: The Non-Negotiable Foundation
An agent that works in the demo and fails in production fails silently unless you can replay its decisions. Tracing is the foundation everything else builds on: every model call with the exact prompt sent and response received, every tool call with arguments and results, every retry and error — organised as a tree, because agent runs are trees. A run spans steps; steps span model calls and tool executions; subagents hang off their orchestrator. Aggregate metrics tell you something regressed; only the trace tells you why — which tool result derailed the plan, which prompt revision changed tool-selection behaviour, where the seventeen-step run went from sensible to absurd. The discipline that pays off most: capture the prompt as actually assembled — after templating, truncation, and context management — not the template you think you sent. The gap between those two is where a large share of production bugs live.
- Trace the tree: run → steps → model/tool calls, with subagent spans nested under parents
- Record exact assembled prompts and raw tool results — reconstructions lie
- Attach token counts, latency, and cost to every span, not just the run total
- OpenTelemetry-based GenAI conventions and LLM-native tracing platforms both work — pick one before launch, not after the incident
Non-Determinism and Regression Testing
Traditional testing assumes the same input yields the same output. Agents break the assumption twice: sampling makes single runs stochastic, and multi-step tasks compound small variations into divergent trajectories. Temperature zero does not rescue you — providers do not guarantee bitwise determinism, and any prompt or model change legitimately reshuffles behaviour. The response is statistical: run each scenario multiple times and assert on the distribution — pass rate at N runs — rather than on one lucky transcript. Assert on outcomes and invariants, not exact wording: the refund was issued, the file compiles, no unauthorised tool was called. Every production failure becomes a permanent regression case. And watch pass-rate trends, not just binary status — a scenario drifting from 95% to 70% is a regression even though it still "passes sometimes."
- Run scenarios N times; assert pass rates and thresholds, not single transcripts
- Test outcomes and invariants — exact-output assertions are flaky by construction
- Convert incidents into regression scenarios: the suite should encode your scar tissue
- Gate deploys on eval results the way you gate on unit tests — prompts are code
Budgets and Failure Taxonomies
An agent loop with no budget is an unbounded liability: a stuck agent can retry, re-read, and re-reason its way through startling token bills. Production agents need hard ceilings enforced by the runtime — maximum steps, maximum tokens, maximum wall-clock, maximum spend per run — with graceful degradation at the limit: summarise progress, save state, escalate; never silently truncate. Latency budgets shape architecture the same way: an interactive product cannot hide a twelve-step sequential loop, which pushes you toward streaming progress, parallel tool calls, and smaller models for routine steps. Equally important is classifying failures when they occur, because the fixes differ. A taxonomy that works in practice: reasoning failures (bad plan), tool failures (external system errors), grounding failures (acting on hallucinated state), context failures (lost or truncated information), and termination failures (loops, premature exits). Tag every failed trace — the histogram tells you where to invest.
- Hard runtime caps on steps, tokens, time, and spend — model self-restraint is not a control
- Degrade gracefully at limits: checkpoint and escalate, never vanish mid-task
- Tag failures by class — reasoning, tool, grounding, context, termination — and fix the biggest bucket first
- Distinguish transient tool errors (retry with backoff) from systematic ones (retrying burns money to fail slower)
Human-in-the-Loop as Architecture
Human oversight fails when it is bolted on as a confirmation dialog; it works when it is designed as part of the control flow. The mechanism that generalises is the hook or gate: defined checkpoints — typically before consequential tool executions — where the runtime pauses, checkpoints state, and routes a decision to a person with enough context to actually decide: what the agent intends, why, and what it would touch. Three design rules carry most of the weight. Gate by consequence, so approvals stay rare enough to be read — a human who approves forty actions an hour is a rubber stamp, and attackers know it. Make interruption first-class: pause, resume, redirect, and abort are states in your run model, not exceptions. And build review capacity for the middle of the autonomy ramp — sampled after-the-fact review of completed runs is how trust is earned before gates are widened, and how drift is caught after.
- Hooks at tool-execution boundaries are the natural gate points — pause, decide, resume
- Approval UX is a security control: intent, rationale, and blast radius on one screen
- Escalation is a feature: an agent that says "I am stuck, here is my state" beats one that improvises
- Sampled review of autonomous runs closes the loop between production and your eval suite
- Agent Engineering gives this two modules — reconstructable traces, graceful degradation, gate design, and the trust curve for widening autonomy
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.