Degrading Gracefully
Every Run Should Have an Exit That Is Not Success
Unattended systems must have a defined behaviour for every way a run can fail to complete, and the default of raising an exception is inadequate because it discards the work and tells the caller nothing actionable. Enumerate the non-success exits and design each: budget exhausted, blocked on a missing precondition, blocked on a permission or an approval, a dependency unavailable, unrecoverably confused. For each one the run should checkpoint its state, produce whatever partial output is legitimately usable, record a specific reason, and hand back a status the caller can branch on. The difference between a system that is operable and one that is not is largely this: whether a person receiving a failed run knows what happened and what to do, or receives a stack trace and a question.
- Enumerate the non-success exits and design behaviour for each one
- Budget exhausted, blocked, unauthorised, dependency down, unrecoverably confused
- Checkpoint, emit usable partial work, record a specific reason, return an actionable status
- Operability is largely whether the recipient of a failure knows what to do next
Partial Results Beat Nothing, With Honest Labelling
An agent that processed sixty of a hundred records has produced real value, and discarding it because the run did not finish wastes both the work and the money. Return the partial result — but label it precisely, because an unlabelled partial result is worse than no result at all: it gets used as though it were complete, and the sixty processed records become a data quality incident that surfaces weeks later. Say what was covered and what was not, in a form a consuming system can act on rather than a sentence a human might read. Make the partial state resumable so continuing does not mean redoing. And decide per feature whether partial output is acceptable at all, because for some tasks the honest answer is that half an answer is not a smaller answer but a wrong one.
- Partial work is real value; discarding it wastes the work and the spend
- An unlabelled partial result gets consumed as complete and becomes a later incident
- State coverage in a form a consuming system can act on, not just readable prose
- For some tasks half an answer is wrong rather than small — decide per feature
Fallbacks and Their Failure Modes
The standard fallbacks are worth having and worth bounding. Retry the run, which works for genuine transient failures and burns money when the failure is systematic — cap it and distinguish the classes before retrying. Fall back to a smaller or different model when a provider is degraded, which requires that your prompts and tool schemas actually work on more than one model family, a portability claim most teams have never tested. Fall back to a deterministic path where one exists, which is another argument for the hybrid architecture. And fall back to a human, which is the most reliable option and needs the handover designed rather than improvised. The failure mode common to all of them is the silent fallback: a system that quietly degrades to a worse path and reports success, so nobody notices until quality drifts. Fallbacks should always be visible in the result and in the metrics.
- Retry only after classifying the failure; systematic failures fail slower and cost more
- Model fallback requires portability you have actually tested, not assumed
- A deterministic fallback path is another reason to keep one around
- Never fall back silently — record it in the result and in the metrics
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.