The AI Learning Hub Journal

What to Log and What Never To

The trace is a copy of everything the model sawthe assembled context includes retrieved documents, user data and tool results — handle the trace store like the data itselfA FULL-FIDELITY TRACE STORE IS A COPY OF A LARGE SLICE OF YOUR SENSITIVE DATAprovisioned as observability infrastructure — give it the access controls and retention of the data, not of a log storethe set of people who can read production traces is defined, not incidentalLOG — THE NON-NEGOTIABLE SETrun and step identifiers, propagated everywherethe assembled prompt, per model calltool arguments as sent, and raw resultsmodel version and prompt version in playbudget state at each stepgate and policy decisions — allows as well as denialsthe final status, with its specific reasonstructured fields, not prose — aggregate questionsneed queries, not parsersNEVER IN THE LOG — AND NOT BY MEMORYcredentials, tokens and keys are scrubbedstructurally in the logging layer, not byremembering not to log themWRITE-TIME REDACTIONnever written, never exposedREAD-TIME FILTERINGa filter someone can bypassthe redaction that matters is the oneapplied before storageMETRICS YOU CAN ANSWER FROM — ONE SCREEN, LOOKED AT WEEKLYsuccess rate by task typecost per success, as a distributionsteps per success · tail latencyfailures counted by classper-tool volume and error ratesargument error rate per toolbudget trips and escalationscache hit rate — drifts silentlyWRITE FOR THE PERSON WHO DID NOT BUILD IT, READING AT THREE IN THE MORNINGunder-logging produces incidents nobody can explain; over-retention copies data nobody should be holding
Log the full non-negotiable set in structured fields, treat the trace store as sensitive data, and scrub secrets at write time

A Trace Contains Everything the Model Saw

This is the fact that changes how traces should be handled. The assembled context includes retrieved documents, user data, tool results and anything the agent read, so a full-fidelity trace store is a copy of a large slice of your sensitive data in a system that was probably provisioned as observability infrastructure with observability access controls. Treat it accordingly: access control it as you would the underlying data, set retention deliberately rather than inheriting a default, and make sure the people who can read production traces are a defined set. Redact at write time rather than at read time, because a redaction applied on the way out is a filter someone can bypass, whereas a value never written is never exposed. Credentials, tokens and keys should be scrubbed structurally in the logging layer, not by remembering not to log them.

  • The trace is a copy of everything the model saw, including retrieved user data
  • Apply the access controls and retention of the underlying data, not of a log store
  • Redact at write time; read-time filtering is a control that can be bypassed
  • Scrub credentials structurally in the logging layer, not by discipline

The Fields That Make Debugging Possible

Against that, log enough to work with, because under-logging produces incidents nobody can explain and a team that guesses. The non-negotiable set: run and step identifiers propagated everywhere, the assembled prompt per model call, tool arguments and raw results, the model and prompt versions, budget state, which policy or gate decisions were made and what they returned, and the final status with its reason. Log allows as well as denials, because knowing that a check ran and passed is evidence and its absence is ambiguous. Structure the fields rather than emitting prose, so the log is queryable and the aggregate questions — which tool has the highest argument error rate this week — are answerable without writing a parser. Emit at a level of detail that a person who did not build the system could follow, since that is who reads it at three in the morning.

  • Identifiers, assembled prompts, arguments, raw results, versions, budget, decisions, final status
  • Record allows as well as denials — an absent record is ambiguous evidence
  • Structured fields, not prose, so aggregate questions are queryable
  • Write for someone who did not build the system and is reading it under pressure

Metrics You Should Be Able to Answer From

Traces explain a run and metrics tell you which runs to look at, so keep a small set that is genuinely used. Success rate by task type. Cost and steps per successful task, reported as a distribution rather than a mean. Tail latency. Failure counts by taxonomy class, which is what turns the previous lesson from vocabulary into a work queue. Per-tool call volume, error rate and argument error rate. Budget-trip rate and escalation rate. Cache hit rate, since it moves cost and latency together and drifts silently. Keep it small enough that the whole set is on one screen and someone looks at it weekly, because the failure mode here is not too few metrics but a dashboard rich enough that nobody reads it and no one notices that a number has been wrong for a month.

  • Success rate by task type; cost and steps per success as distributions, not means
  • Failure counts by taxonomy class — that is what makes the taxonomy operational
  • Per-tool volume, error rate and argument error rate; budget trips and escalations
  • Keep the set small enough that it fits on a screen and gets read weekly

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.