The AI Learning Hub Journal

Testing Agents That Take Several Steps

The error shows up at the end — the cause lives at the startan agent run has an outcome and a trajectory, and grading only one of them misleads in a different way eachPER-STEP ASSERTIONS — FAIL WHERE THE PREMISE BROKEstep 1 — parses the IDsubtly wrongstep 2 — looks it upconfident, on a bad premisestep 3 — updates a recordthe wrong onevisible failurecause: step 1assert here: valid ID?→ state what must be true after each step — the run now fails at step 1, not step 3and check the tool calls themselves: right tool, valid arguments, arguments matching what the user actually askedOUTCOME-ONLY GRADINGpasses the lucky run — right end state reachedthrough six wasteful calls and one destructiveaction that happened to be reversiblean incident waiting for the luck to run outTRAJECTORY-ONLY GRADINGfails the agent that found a better path thanthe route you imagined when writing the testpunishes improvement for being unfamiliarGATE ON THE OUTCOME — READ THE TRAJECTORY AS DIAGNOSIS, NOT VERDICTWORK DONE IS A QUALITY SIGNALtrack steps, tool calls, tokens and cost persuccessful run — a big rise is a regressioneven when the success rate improvedread the tail — looping and retrying live thereCAPTURE ENOUGH TO REPLAYassembled prompts, exact arguments, rawresults, model versions, starting state —then fork the trace at the failing stepa failure you cannot replay never becomes a caseGATE ON OUTCOME, ASSERT PER STEP, AND KEEP EVERY REPLAYABLE FAILUREan agent eval suite is largely a collection of past failures with assertions attached
Gate agent runs on the final outcome, but assert after every step — late failures have early causes, and only replayable failures become regression cases.

Outcome and Trajectory Are Different Questions

An agent run has two things worth grading and teams routinely grade only one. The outcome is whether the final state of the world is what it should be — the record updated, the file written, the ticket in the right queue — and it is what the user cares about. The trajectory is how the run got there: which tools were called, in what order, with what arguments, how many steps it took, what it read along the way. Outcome-only grading passes a run that reached the right end state through six unnecessary calls, two failed attempts and a destructive action that happened to be reversible, and that run is an incident waiting for the case where the luck runs out. Trajectory-only grading fails in the other direction, penalising an agent for finding a better path than the one you imagined. Gate on the outcome and read the trajectory to understand why the outcome happened.

  • Outcome: is the final world state correct? Trajectory: how did it get there?
  • Outcome-only grading passes lucky runs that took wasteful or dangerous paths
  • Trajectory-only grading punishes agents for improving on your imagined route
  • Gate on outcome; use trajectory as the diagnosis, not the verdict

Failures That Surface Late

The hardest agent failures to diagnose are the ones where step three fails because step one was subtly wrong — a mis-parsed identifier, a date read in the wrong format, a search that returned plausible but incorrect records — and every step afterwards operates confidently on a bad premise. The visible error is at the end and the cause is at the beginning, so a grader that only inspects the final output tells you almost nothing. The counter is per-step assertions: state what must be true after each step and check it there, so the run fails at the point the premise broke rather than at the point the consequence became visible. Tool-call correctness belongs in the same layer — the right tool, valid arguments, and arguments that match what the user actually asked for. This costs real work to write and it is the highest-value thing you can add to an agent eval.

  • Late failures usually have early causes; the visible error is not the defect
  • Assert post-conditions after each step so the run fails where the premise broke
  • Check tool-call correctness: right tool, valid arguments, arguments matching intent
  • Per-step assertions turn an afternoon of trace reading into a failure that names itself

Cost and Step Count Are Quality Signals

In a multi-step system, how much work was done is information about how well the work was understood. A run that reaches the correct outcome in fifteen steps where four would do is telling you something about the plan, the tool descriptions, or the clarity of the task definition — and it is also telling you what the feature will cost at volume. Track steps, tool calls, tokens, wall-clock time and cost per successful run, and treat a significant increase as a regression needing an explanation even when the success rate improved. Watch the shape of the distribution rather than the mean, because the runs in the tail are usually the ones looping, retrying the same failing call, or wandering, and they are simultaneously the expensive cases and the most diagnostic ones. A step budget that terminates a run is a guardrail rather than a fix; the runs that hit it are your reading list.

  • Steps and cost per successful run belong in the same report as quality
  • A large increase in work is a regression even when the success rate rose
  • Read the tail — looping, retrying and wandering runs all live there
  • Step budgets contain the damage; the runs that hit them are the diagnosis

Replaying a Failed Run

The value of a failed run lies entirely in whether you can get back to it. That requires the trace to record the assembled prompt at each step rather than the template, the exact tool arguments and the raw results returned, the model and prompt versions in play, and the state the environment started in. Reconstructing any of those from memory produces a plausible story rather than a cause. With them captured, three things become possible: replay the run against a fix and confirm it now passes, replay it against a different model family to see whether the failure is model-specific, and fork the trace at the step that went wrong so you can test a change without re-running everything before it. Then promote the case into the suite. An agent eval suite is largely a collection of past failures with assertions attached, and the ones nobody could reconstruct never made it in.

  • Capture assembled prompts, exact arguments, raw results, versions, and initial state
  • Replay against a fix, and against another model family, to isolate the cause
  • Fork the trace at the failing step instead of re-running the whole trajectory
  • A failure you cannot reproduce can never become a regression case

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.