Tracing a Reconstructable Run
The Run Is a Tree
Agent execution is not a linear log, and forcing it into one is why so many teams have logs they cannot use. A run contains steps; a step contains one model call and the tool executions it requested; a subagent hangs beneath the step that spawned it; a retry sits under the call it repeats. Model it as spans with parent references and a single run identifier propagated everywhere, including into downstream systems the tools touch, so an anomalous record in another service can be traced back to the run and step that created it. The structure is what makes the trace usable: it lets you collapse a forty-step run to its shape, expand only the step that went wrong, and see immediately whether a subagent was involved. A flat stream of log lines from a concurrent agent is close to unreadable and reliably goes unread.
- Run contains steps; steps contain a model call and its tool executions
- Subagents nest under the spawning step; retries nest under the call they repeat
- Propagate the run identifier into downstream systems the tools write to
- Structure is what makes a long run readable — flat logs from concurrent agents do not get read
Record What Cannot Be Reconstructed
The test for whether a field belongs in the trace is whether you could recover it afterwards from something else. Assembled prompts cannot be recovered, because templating, retrieval, compaction and truncation all happened at runtime and the template does not tell you what came out — this is the single most frequently missing field and the one whose absence most often ends an investigation. Raw tool results cannot be recovered, because the external system has moved on. Tool arguments as sent, versions in play, budget state at each step, and which branch of your assembly logic ran all fall in the same category. What you can safely omit is anything derivable: token counts you can recompute, formatting you can regenerate. Store the inputs to decisions rather than the narrative of them.
- Assembled prompts are unrecoverable afterwards and are the field most often missing
- Raw tool results are unrecoverable because the external system has moved on
- Also capture arguments as sent, versions, budget state, and which assembly branch ran
- Omit anything you can derive later; store the inputs to decisions, not the narrative
Volume, Sampling and Cost
Full-fidelity traces are large — they contain every prompt, which means they can exceed the size of the work product by a wide margin — and at production volume the storage and the transmission are a real cost that teams discover late and respond to by turning tracing down to nothing. Plan the tiering instead. Keep full fidelity for a sampled fraction of successful runs, and for every run that failed, hit a budget ceiling, triggered a gate, or ended in escalation, since those are the ones anyone will want. Keep skeletons — step count, tools called, timings, costs, outcome — for everything, because the skeleton is what aggregate analysis runs on and it is small. Set retention by class rather than uniformly, and decide it before launch, because after an incident the retention window you had is the retention window you have.
- Full traces can exceed the size of the work product; volume costs are real
- Full fidelity for failures, budget trips, gated and escalated runs, plus a success sample
- Skeletons for everything: steps, tools, timings, cost, outcome
- Set retention per class before launch — after an incident it is too late
Try It Yourself
The test of a trace is not how much it holds but whether someone can rebuild a run from it after the fact. Run that test on a real failure now, rather than during the incident where it matters.
Find one run of your agent that failed or behaved strangely and try to answer the nine questions below from the stored trace alone - no access to the running system, no re-execution, no asking the engineer who built it. Every question you cannot answer names a field to add. If you have no traces yet, run the list against the trace schema you are designing and mark which fields it would already carry.
Run: Outcome: Answer from the trace alone: 1. What exact context did the model see at the step where it went wrong? 2. What arguments were sent to each tool, as sent rather than as intended? 3. What did each tool return, raw? 4. Which model, prompt, tool schema and retrieval index versions were in force? 5. What was the remaining budget at each step? 6. Which branch of the assembly logic ran, and was anything truncated or compacted? 7. Did a subagent run, and under which step? 8. Which gate or policy decisions were evaluated, and what did each return? 9. Why did the run end, as recorded rather than as inferred? For each question you could not answer: Field to add | Emitted where in the code | Recoverable from anything else? | Retention class
- You reached a specific verdict on the cause, or you can name the exact question the trace could not answer
- Every gap became a named field with a place in the emitting code, not a note to improve logging
- You can say which gaps were genuinely unrecoverable and which you could have derived from what was already stored
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.