The Failure Taxonomy You Will Actually Meet
Wrong Tool, and Right Tool Wrong Arguments
These are the two most common failures in production and they need separating, because the fixes differ and the aggregate metric hides both. Wrong tool means the model selected an action inappropriate to the situation, and it points at the toolset: overlapping capabilities, a description that omits when not to use this, or a missing tool that made a near-fit look reasonable. Right tool wrong arguments means selection was correct and the invocation was not — a wrong identifier, a misparsed date, a free-text field filled with an invented value, a unit mismatch — and it points at the schema. Instrument both separately: per-tool selection rate against a labelled expectation, and per-tool argument validation failure rate. When one tool dominates either metric, you have found your next hour of work, and it is usually a rewrite of one description.
- Wrong tool indicts the toolset: overlap, missing "when not to", or an absent capability
- Wrong arguments indict the schema: loose types, unstated formats, undescribed fields
- Measure selection accuracy and argument-error rate per tool, separately
- One tool usually dominates, and the fix is usually one description
Hallucinated Tools and Malformed Calls
Sometimes a model calls something that does not exist, or invents a parameter, or produces a call that does not conform to the schema. The instinct is to treat this as a model defect, and it is more usefully treated as a signal about your context. Invented tools usually correspond to a capability the agent needed and did not have, so the name it invented is a specification for the tool you are missing, and reading those names in aggregate is one of the cheapest product research exercises available. Invented parameters usually mean the schema does not express something the task requires. Operationally, handle these without drama: validate every call against the registered schema before dispatch, return a specific correctable error naming what is available, and count them, because a rising rate normally follows a change you made rather than anything the model did.
- An invented tool name is a specification for the capability you did not provide
- An invented parameter usually means the schema cannot express the task
- Validate against the registered schema before dispatch and return a correctable error
- Track the rate: it usually rises after your change, not after a model change
Premature Stop and Infinite Loop
The two termination failures are opposites with a shared cause, which is a completion decision left to judgement. Premature stop is the expensive one because it is quiet: the agent declares success on partial work and the failure surfaces later, downstream, as a data problem that nobody attributes to the agent. Its signatures are a low step count with a successful status, a terminal call with thin or missing evidence fields, and success rates that improve suspiciously after a prompt change that emphasised efficiency. Infinite loop is loud, costs money and trips a budget, and its signatures are repeated action fingerprints and no movement in the progress predicate. Both are addressed by the same discipline covered earlier: a deterministic completion predicate where one exists, a terminal tool carrying evidence, and stuck-detection running inside the loop rather than in a dashboard afterwards.
- Premature stop is quiet and surfaces downstream as a data problem nobody attributes
- Signatures: low step count with success, thin evidence, suspicious efficiency gains
- Infinite loop is loud: repeated fingerprints and a flat progress predicate
- Both are fixed by checkable completion, evidence-carrying terminals and in-loop detection
Context Exhaustion and Compounding Premise Errors
Context exhaustion is the failure that looks like a model getting worse mid-run: quality degrades as the window fills, the agent forgets a constraint agreed twenty steps ago, and eventually something overflows or is silently trimmed. Its fix is upstream in context engineering rather than here, but its detection belongs in operations — track context size per step and correlate failures against it, and you will usually find a threshold beyond which quality falls off. The subtler relative is the compounding premise error: step two read the wrong record, and every subsequent step operated competently on a false premise, so the visible error is at the end and the cause is near the beginning. The operational counter is checkpoint assertions after the steps that establish premises, so the run fails where the premise broke instead of where the consequence became visible.
- Track context size per step and correlate against failures to find the threshold
- Degradation before overflow is the common case; the hard limit is not the real limit
- Compounding premise errors put the visible error far from the cause
- Assert after premise-establishing steps so the run fails where it actually broke
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.