The AI Learning Hub Journal

Observability and Closing the Loop

Closing the loop between what ships and what you testthe loop only counts as closed when real failures end up back in the suitePRODUCTION IS THE BESTSOURCE OF EVAL CASESReal users generate failure modesthat nobody on the team thoughtto write down. Every one you findis a test case you did not haveto invent — take it and keep it.1Traced production runsinputs, tool calls, outputs2Failures surfacedclustered and prioritised3Added to the eval setthe representative ones4Fix the causeprompt, retrieval or guard5Regression runthe whole set, not the fix6Redeploythen watch the traces againWithout the trace you never see the failure; without the eval case you fix it once and lose it
An eval set that never grows is a set that stops catching things — production is where the next test case comes from

Traces Are the Substrate

Offline evals tell you how the system performs on cases you have collected; production tells you what is actually happening, and the connection between them is tracing. An agent run is a tree — run, steps, model calls, tool calls, subagent spans — and the trace must capture the prompt as assembled after templating, truncation, and context management rather than the template you believe you sent, because the gap between those two is where a large share of production bugs live. Attach tokens, latency, and cost to every span, along with the identifiers needed to slice later. Whether you use OpenTelemetry's GenAI conventions or an LLM-native tracing platform matters far less than picking one before launch, because retrofitting tracing during an incident is a uniquely bad experience.

  • Trace the tree: run, steps, model and tool calls, nested subagent spans
  • Capture the assembled prompt and raw tool results — reconstructions mislead
  • Attach cost, tokens, and latency per span, not just per run
  • Choose the tracing approach before launch; retrofitting during an incident is painful

Online Signals and Guardrail Evals

Production carries quality signals if you instrument for them. Implicit ones are the most abundant and the most honest: retries, rephrasings of the same question, conversation abandonment, escalation to a human, users copying only part of an answer, edits made to generated content before it is used. Explicit feedback is sparse and biased toward extremes, useful mainly as a pointer to traces worth reading. On top of these, run lightweight automated checks over live traffic — schema validity, groundedness against retrieved sources, refusal rate, policy checks — sampled at a rate you can afford, giving a continuous quality signal between offline runs. Alert on movement in those signals rather than only on errors, because the characteristic production failure here is a quality slide with a perfectly healthy error rate.

  • Implicit signals — retry, rephrase, abandon, escalate, edit — outnumber and outperform thumbs
  • Sampled online checks give continuous quality signal between offline runs
  • Alert on quality drift, not just errors; degradation is silent by default
  • Segment online metrics the same way you slice offline evals

The Loop That Closes

The point of production observability is not the dashboard, it is the pipeline from failure to permanent test. Make it a defined workflow with an owner: someone reviews flagged and sampled traces on a fixed cadence, classifies each failure by mode, and converts the informative ones into eval cases with expected outcomes and graders. New cases enter the dev set, get confirmed as currently failing, and stay in the suite forever after the fix. Two habits keep this alive. First, define a target for how quickly a production failure becomes a case — an incident that has not produced a regression test within the week usually never will. Second, track the failure-mode distribution over time, because the shape of that histogram is what tells you whether the last quarter of work actually moved anything.

  • Named owner and fixed cadence for trace review — otherwise it does not happen
  • Every incident becomes a permanent case; the suite is institutional memory
  • Set a time target from production failure to eval case
  • Track the failure-mode distribution over time as the real progress metric

Guarding the Gap Between Offline and Online

Offline scores and production quality drift apart, and the divergence is itself a metric worth watching. Common causes: the input distribution has moved and the eval set no longer looks like traffic, the eval environment has diverged from production configuration, or the suite has been tuned against until it flatters the system. Check the relationship deliberately — sample recent production traces, grade them with the same rubrics you use offline, and compare against what the suite predicts. A persistent gap means the suite has stopped being representative and needs refreshing from current traffic rather than defending. This check is cheap, most teams never run it, and it is the difference between a suite that reflects the product and a suite that reflects the product as it was when someone last cared.

  • Grade sampled production traffic with the offline rubrics and compare
  • A persistent offline-online gap means the case mix has gone stale
  • Verify the eval environment matches production configuration, not an idealised version
  • Refresh from current traffic on a schedule, not when someone complains

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.