The AI Learning Hub Journal

Degrading Gracefully

Every run needs an exit that is not successenumerate the ways a run can fail to complete and design each one — an exception discards the work and explains nothingTHE NON-SUCCESS EXITS — EACH ONE A DESIGNED BEHAVIOURbudget exhaustedmissing preconditionpermission / approvaldependency unavailableunrecoverably confusedeach exit checkpoints state · emits usable partial output · records a specific reason · returns a branchable statusthe caller can branch on the status without reading a stack traceFALL DOWN A LADDER, NOT OFF A CLIFF — EACH RUNG BOUNDED AND VISIBLE1 · RETRY THE RUNtransients only — classify the failure first, cap the attempts2 · FALL BACK TO ANOTHER MODELneeds portability you have actually tested, not assumed3 · FALL BACK TO THE DETERMINISTIC PATHanother reason the hybrid design keeps one around4 · HAND OVER TO A HUMANthe most reliable option — design the handoverTHE SILENT FALLBACK — THE SHARED FAILURE MODEthe system quietly degrades to a worse pathand reports ordinary success — nobodynotices until quality has drifted for weeksrecord every fallback in the result and the metricsTHE OPERABILITY TESTthe person receiving a failed run knows whathappened and what to do next — not a stacktrace and a questionPARTIAL RESULTS BEAT NOTHING — WITH HONEST, MACHINE-READABLE LABELLINGan unlabelled sixty-of-a-hundred is consumed as complete and becomes a later data incident — state coverage, make it resumableAN OPERABLE SYSTEM FAILS WITH A REASON, PARTIAL OUTPUT AND A NEXT STEPand for some tasks half an answer is a wrong answer, not a smaller one — decide that per feature, in advance
Enumerate every non-success exit and fall down a bounded, visible ladder of fallbacks — a silent fallback reports success while quality drifts

Every Run Should Have an Exit That Is Not Success

Unattended systems must have a defined behaviour for every way a run can fail to complete, and the default of raising an exception is inadequate because it discards the work and tells the caller nothing actionable. Enumerate the non-success exits and design each: budget exhausted, blocked on a missing precondition, blocked on a permission or an approval, a dependency unavailable, unrecoverably confused. For each one the run should checkpoint its state, produce whatever partial output is legitimately usable, record a specific reason, and hand back a status the caller can branch on. The difference between a system that is operable and one that is not is largely this: whether a person receiving a failed run knows what happened and what to do, or receives a stack trace and a question.

  • Enumerate the non-success exits and design behaviour for each one
  • Budget exhausted, blocked, unauthorised, dependency down, unrecoverably confused
  • Checkpoint, emit usable partial work, record a specific reason, return an actionable status
  • Operability is largely whether the recipient of a failure knows what to do next

Partial Results Beat Nothing, With Honest Labelling

An agent that processed sixty of a hundred records has produced real value, and discarding it because the run did not finish wastes both the work and the money. Return the partial result — but label it precisely, because an unlabelled partial result is worse than no result at all: it gets used as though it were complete, and the sixty processed records become a data quality incident that surfaces weeks later. Say what was covered and what was not, in a form a consuming system can act on rather than a sentence a human might read. Make the partial state resumable so continuing does not mean redoing. And decide per feature whether partial output is acceptable at all, because for some tasks the honest answer is that half an answer is not a smaller answer but a wrong one.

  • Partial work is real value; discarding it wastes the work and the spend
  • An unlabelled partial result gets consumed as complete and becomes a later incident
  • State coverage in a form a consuming system can act on, not just readable prose
  • For some tasks half an answer is wrong rather than small — decide per feature

Fallbacks and Their Failure Modes

The standard fallbacks are worth having and worth bounding. Retry the run, which works for genuine transient failures and burns money when the failure is systematic — cap it and distinguish the classes before retrying. Fall back to a smaller or different model when a provider is degraded, which requires that your prompts and tool schemas actually work on more than one model family, a portability claim most teams have never tested. Fall back to a deterministic path where one exists, which is another argument for the hybrid architecture. And fall back to a human, which is the most reliable option and needs the handover designed rather than improvised. The failure mode common to all of them is the silent fallback: a system that quietly degrades to a worse path and reports success, so nobody notices until quality drifts. Fallbacks should always be visible in the result and in the metrics.

  • Retry only after classifying the failure; systematic failures fail slower and cost more
  • Model fallback requires portability you have actually tested, not assumed
  • A deterministic fallback path is another reason to keep one around
  • Never fall back silently — record it in the result and in the metrics

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.