Testing Agents That Take Several Steps
Outcome and Trajectory Are Different Questions
An agent run has two things worth grading and teams routinely grade only one. The outcome is whether the final state of the world is what it should be — the record updated, the file written, the ticket in the right queue — and it is what the user cares about. The trajectory is how the run got there: which tools were called, in what order, with what arguments, how many steps it took, what it read along the way. Outcome-only grading passes a run that reached the right end state through six unnecessary calls, two failed attempts and a destructive action that happened to be reversible, and that run is an incident waiting for the case where the luck runs out. Trajectory-only grading fails in the other direction, penalising an agent for finding a better path than the one you imagined. Gate on the outcome and read the trajectory to understand why the outcome happened.
- Outcome: is the final world state correct? Trajectory: how did it get there?
- Outcome-only grading passes lucky runs that took wasteful or dangerous paths
- Trajectory-only grading punishes agents for improving on your imagined route
- Gate on outcome; use trajectory as the diagnosis, not the verdict
Failures That Surface Late
The hardest agent failures to diagnose are the ones where step three fails because step one was subtly wrong — a mis-parsed identifier, a date read in the wrong format, a search that returned plausible but incorrect records — and every step afterwards operates confidently on a bad premise. The visible error is at the end and the cause is at the beginning, so a grader that only inspects the final output tells you almost nothing. The counter is per-step assertions: state what must be true after each step and check it there, so the run fails at the point the premise broke rather than at the point the consequence became visible. Tool-call correctness belongs in the same layer — the right tool, valid arguments, and arguments that match what the user actually asked for. This costs real work to write and it is the highest-value thing you can add to an agent eval.
- Late failures usually have early causes; the visible error is not the defect
- Assert post-conditions after each step so the run fails where the premise broke
- Check tool-call correctness: right tool, valid arguments, arguments matching intent
- Per-step assertions turn an afternoon of trace reading into a failure that names itself
Cost and Step Count Are Quality Signals
In a multi-step system, how much work was done is information about how well the work was understood. A run that reaches the correct outcome in fifteen steps where four would do is telling you something about the plan, the tool descriptions, or the clarity of the task definition — and it is also telling you what the feature will cost at volume. Track steps, tool calls, tokens, wall-clock time and cost per successful run, and treat a significant increase as a regression needing an explanation even when the success rate improved. Watch the shape of the distribution rather than the mean, because the runs in the tail are usually the ones looping, retrying the same failing call, or wandering, and they are simultaneously the expensive cases and the most diagnostic ones. A step budget that terminates a run is a guardrail rather than a fix; the runs that hit it are your reading list.
- Steps and cost per successful run belong in the same report as quality
- A large increase in work is a regression even when the success rate rose
- Read the tail — looping, retrying and wandering runs all live there
- Step budgets contain the damage; the runs that hit them are the diagnosis
Replaying a Failed Run
The value of a failed run lies entirely in whether you can get back to it. That requires the trace to record the assembled prompt at each step rather than the template, the exact tool arguments and the raw results returned, the model and prompt versions in play, and the state the environment started in. Reconstructing any of those from memory produces a plausible story rather than a cause. With them captured, three things become possible: replay the run against a fix and confirm it now passes, replay it against a different model family to see whether the failure is model-specific, and fork the trace at the step that went wrong so you can test a change without re-running everything before it. Then promote the case into the suite. An agent eval suite is largely a collection of past failures with assertions attached, and the ones nobody could reconstruct never made it in.
- Capture assembled prompts, exact arguments, raw results, versions, and initial state
- Replay against a fix, and against another model family, to isolate the cause
- Fork the trace at the failing step instead of re-running the whole trajectory
- A failure you cannot reproduce can never become a regression case
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.