The AI Learning Hub Journal

The Failure Taxonomy You Will Actually Meet

Eight failures, four families — each points at a different fixclassify before you fix: each class indicts a specific part of the system, and the metric that finds it differsTHE TWO MOST COMMON — SELECTION AND ARGUMENTSWRONG TOOL — INDICTS THE TOOLSEToverlapping capabilities, a missing 'whennot to use this', or an absent toolRIGHT TOOL, WRONG ARGUMENTS — THE SCHEMAloose types, unstated formats and units,undescribed fields invite plausible wrong valuesone tool usually dominates the metric — the fix is one descriptionINVENTED CALLS — SIGNALS ABOUT YOUR CONTEXTHALLUCINATED TOOL — A SPECIFICATIONthe invented name describes the capabilitythe agent needed and was not givenINVENTED PARAMETER — A SCHEMA GAPthe schema cannot express somethingthe task genuinely requiresvalidate against the registry before dispatch, and count themTERMINATION — COMPLETION LEFT TO JUDGEMENTPREMATURE STOP — THE QUIET ONEsuccess declared on partial work — surfacesdownstream as a data problem, unattributedINFINITE LOOP — THE LOUD ONErepeated action fingerprints, a flatprogress predicate, a tripped budgetcheckable completion, evidence-carrying terminals, in-loop detectionCONTEXT — THE CAUSE FAR FROM THE SYMPTOMCONTEXT EXHAUSTION — MID-RUN DECAYquality falls as the window fills — trackcontext size per step; find the thresholdCOMPOUNDING PREMISE ERRORstep two read the wrong record; every laterstep worked competently on a false premiseassert after premise-establishing steps, so the run fails where it brokeCOUNT FAILURES BY CLASS — THAT IS WHAT TURNS A TAXONOMY INTO A WORK QUEUEthe aggregate failure rate hides all of this; the class distribution tells you what to fix next
Classify failures before fixing them — each class indicts a specific component and carries its own metric

Wrong Tool, and Right Tool Wrong Arguments

These are the two most common failures in production and they need separating, because the fixes differ and the aggregate metric hides both. Wrong tool means the model selected an action inappropriate to the situation, and it points at the toolset: overlapping capabilities, a description that omits when not to use this, or a missing tool that made a near-fit look reasonable. Right tool wrong arguments means selection was correct and the invocation was not — a wrong identifier, a misparsed date, a free-text field filled with an invented value, a unit mismatch — and it points at the schema. Instrument both separately: per-tool selection rate against a labelled expectation, and per-tool argument validation failure rate. When one tool dominates either metric, you have found your next hour of work, and it is usually a rewrite of one description.

  • Wrong tool indicts the toolset: overlap, missing "when not to", or an absent capability
  • Wrong arguments indict the schema: loose types, unstated formats, undescribed fields
  • Measure selection accuracy and argument-error rate per tool, separately
  • One tool usually dominates, and the fix is usually one description

Hallucinated Tools and Malformed Calls

Sometimes a model calls something that does not exist, or invents a parameter, or produces a call that does not conform to the schema. The instinct is to treat this as a model defect, and it is more usefully treated as a signal about your context. Invented tools usually correspond to a capability the agent needed and did not have, so the name it invented is a specification for the tool you are missing, and reading those names in aggregate is one of the cheapest product research exercises available. Invented parameters usually mean the schema does not express something the task requires. Operationally, handle these without drama: validate every call against the registered schema before dispatch, return a specific correctable error naming what is available, and count them, because a rising rate normally follows a change you made rather than anything the model did.

  • An invented tool name is a specification for the capability you did not provide
  • An invented parameter usually means the schema cannot express the task
  • Validate against the registered schema before dispatch and return a correctable error
  • Track the rate: it usually rises after your change, not after a model change

Premature Stop and Infinite Loop

The two termination failures are opposites with a shared cause, which is a completion decision left to judgement. Premature stop is the expensive one because it is quiet: the agent declares success on partial work and the failure surfaces later, downstream, as a data problem that nobody attributes to the agent. Its signatures are a low step count with a successful status, a terminal call with thin or missing evidence fields, and success rates that improve suspiciously after a prompt change that emphasised efficiency. Infinite loop is loud, costs money and trips a budget, and its signatures are repeated action fingerprints and no movement in the progress predicate. Both are addressed by the same discipline covered earlier: a deterministic completion predicate where one exists, a terminal tool carrying evidence, and stuck-detection running inside the loop rather than in a dashboard afterwards.

  • Premature stop is quiet and surfaces downstream as a data problem nobody attributes
  • Signatures: low step count with success, thin evidence, suspicious efficiency gains
  • Infinite loop is loud: repeated fingerprints and a flat progress predicate
  • Both are fixed by checkable completion, evidence-carrying terminals and in-loop detection

Context Exhaustion and Compounding Premise Errors

Context exhaustion is the failure that looks like a model getting worse mid-run: quality degrades as the window fills, the agent forgets a constraint agreed twenty steps ago, and eventually something overflows or is silently trimmed. Its fix is upstream in context engineering rather than here, but its detection belongs in operations — track context size per step and correlate failures against it, and you will usually find a threshold beyond which quality falls off. The subtler relative is the compounding premise error: step two read the wrong record, and every subsequent step operated competently on a false premise, so the visible error is at the end and the cause is near the beginning. The operational counter is checkpoint assertions after the steps that establish premises, so the run fails where the premise broke instead of where the consequence became visible.

  • Track context size per step and correlate against failures to find the threshold
  • Degradation before overflow is the common case; the hard limit is not the real limit
  • Compounding premise errors put the visible error far from the cause
  • Assert after premise-establishing steps so the run fails where it actually broke

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.