What to Log and What Never To
A Trace Contains Everything the Model Saw
This is the fact that changes how traces should be handled. The assembled context includes retrieved documents, user data, tool results and anything the agent read, so a full-fidelity trace store is a copy of a large slice of your sensitive data in a system that was probably provisioned as observability infrastructure with observability access controls. Treat it accordingly: access control it as you would the underlying data, set retention deliberately rather than inheriting a default, and make sure the people who can read production traces are a defined set. Redact at write time rather than at read time, because a redaction applied on the way out is a filter someone can bypass, whereas a value never written is never exposed. Credentials, tokens and keys should be scrubbed structurally in the logging layer, not by remembering not to log them.
- The trace is a copy of everything the model saw, including retrieved user data
- Apply the access controls and retention of the underlying data, not of a log store
- Redact at write time; read-time filtering is a control that can be bypassed
- Scrub credentials structurally in the logging layer, not by discipline
The Fields That Make Debugging Possible
Against that, log enough to work with, because under-logging produces incidents nobody can explain and a team that guesses. The non-negotiable set: run and step identifiers propagated everywhere, the assembled prompt per model call, tool arguments and raw results, the model and prompt versions, budget state, which policy or gate decisions were made and what they returned, and the final status with its reason. Log allows as well as denials, because knowing that a check ran and passed is evidence and its absence is ambiguous. Structure the fields rather than emitting prose, so the log is queryable and the aggregate questions — which tool has the highest argument error rate this week — are answerable without writing a parser. Emit at a level of detail that a person who did not build the system could follow, since that is who reads it at three in the morning.
- Identifiers, assembled prompts, arguments, raw results, versions, budget, decisions, final status
- Record allows as well as denials — an absent record is ambiguous evidence
- Structured fields, not prose, so aggregate questions are queryable
- Write for someone who did not build the system and is reading it under pressure
Metrics You Should Be Able to Answer From
Traces explain a run and metrics tell you which runs to look at, so keep a small set that is genuinely used. Success rate by task type. Cost and steps per successful task, reported as a distribution rather than a mean. Tail latency. Failure counts by taxonomy class, which is what turns the previous lesson from vocabulary into a work queue. Per-tool call volume, error rate and argument error rate. Budget-trip rate and escalation rate. Cache hit rate, since it moves cost and latency together and drifts silently. Keep it small enough that the whole set is on one screen and someone looks at it weekly, because the failure mode here is not too few metrics but a dashboard rich enough that nobody reads it and no one notices that a number has been wrong for a month.
- Success rate by task type; cost and steps per success as distributions, not means
- Failure counts by taxonomy class — that is what makes the taxonomy operational
- Per-tool volume, error rate and argument error rate; budget trips and escalations
- Keep the set small enough that it fits on a screen and gets read weekly
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.