Return Values That Enable Recovery
The Return Is the Next Prompt
Whatever a tool returns becomes context, and it is the material the model uses to choose the next action. Designing returns for a human reading a log is the wrong target. Three things belong in almost every return. The outcome, stated explicitly rather than implied by the presence or absence of data, because an empty result and a failed query are different situations that look identical when both render as nothing. The resulting state where the tool changed something, so the model does not have to re-read to find out what happened. And the affordances, meaning what can be done next, particularly identifiers needed for a follow-up call. A return that says the update succeeded and the record now reads as follows saves a verification step on every single call, which compounds across a run.
- The return value is the next prompt — design it for the model, not for a log reader
- State the outcome explicitly; empty results and failures render identically otherwise
- Include resulting state after a mutation so the model need not re-read
- Carry forward the identifiers a follow-up call will need
Errors Should Say What to Do
An error message is an instruction to the model, so write it as one. A stack trace or a generic failure produces a retry of the same call, because there is nothing in the message to act on differently. Say what was wrong, specifically, and say what would be valid — this identifier was not found, and identifiers for this resource are eight digits; or this account is closed, so no further transactions can be posted to it. Distinguish three classes clearly, because the correct behaviour differs: retryable transient conditions, which should say so and indicate a wait; correctable input errors, which should state the correction; and terminal conditions, which should say that no retry will help so the agent stops trying and either takes another route or escalates. Most looping in production traces back to error text that failed to communicate which class applied.
- Write errors as instructions: what was wrong and what would be valid
- Distinguish transient, correctable and terminal explicitly in the message
- Terminal errors must say that retrying will not help, or the agent will retry
- A large share of production looping is error text that omits the class
Volume, Truncation and Pointers
Tools that can return unbounded output are a standing threat to the context budget, and the failure is abrupt: one query returns ten thousand rows and the run is over. Bound every return at the tool, not downstream. Default to a modest page with the total count stated, so the model knows what it is not seeing and can refine rather than assume it has everything. Where content must be truncated, truncate visibly with a marker and a handle for retrieving the rest, since silent truncation produces an agent reasoning confidently over a fragment it believes is whole. Where a large artefact is genuinely needed, write it somewhere addressable and return the reference, so the content enters context only if a later step actually requires it. The general principle: returns should be summaries with retrieval paths, not payload dumps.
- Bound output at the tool; unbounded returns end runs abruptly
- Return a page plus the total count so the model knows what it has not seen
- Truncate visibly with a handle for the rest — silent truncation produces confident fragments
- Write large artefacts to an addressable location and return a reference
Try It Yourself
A tool definition is read once by a consumer that cannot ask what you meant and cannot experiment to find out. Take one of yours and rewrite it for that reader, description through return values.
Pick the tool in your system with the highest argument-error or misselection rate. If you have no traces yet, pick the one whose side effect you would least like performed twice. Rewrite its definition against the template below, including the return value for success and for each of the three error classes. If you own no tools, do this for one tool an agent product exposes to you, written from its observable behaviour, and note what you had to guess.
TOOL: DESCRIPTION - What it does to the world: - When to use it: - When NOT to use it, and which sibling tool to prefer instead: - Preconditions that must hold: - Cost or latency, if it should affect selection: - Example call: SCHEMA - one row per field field | type | required | allowed values (enumerate fixed sets) | format and units | description RETURNS - Success: outcome stated explicitly, resulting state after the mutation, identifiers a follow-up call needs - Empty result: distinguishable from failure, with the total count - Transient error: says it is transient and how long to wait - Correctable error: says what was wrong and what would be valid - Terminal error: says that retrying will not help, and what route remains - Bounds: page size, total count, and the handle for retrieving the rest
- The description names at least one situation where this tool is the wrong choice and says which tool to use there instead
- No field with a fixed set of legal values is typed as a free string
- For each of the three error strings you can state what the model's next action would be after reading it
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.