The AI Learning Hub Journal

Escalation as a First-Class Outcome

Asking for help is an action, not a failurean agent with no way to stop and ask will improvise, because continuing is the only action available to itTHE RUN BLOCKSrather than improvisingIT ESCALATESa schema'd tool callA PERSON DECIDESrouted by decision typeTHE RUN RESUMESfrom its checkpointthe answer resumes the checkpointed run — it never requires starting a fresh oneTHE ESCALATION SCHEMA — WHAT IT SENDSwhat it was trying to dowhat it has established so farwhat is blocking itwhat it would do next, if told to proceedwhat it needs from a personthe schema is what makes asking a legitimate outcomeTHE HANDOVER — A DELIVERABLE, NOT A PINGthe goal, in the user's original wordswhat was done, and the state it leavesthe specific question being askedthe options seen, with the agent's pick and whya link to the full trace, for whoever wants ita bare notification moves the work without the contextTRIGGERS ARE STRUCTURAL FIRST — THE RUNTIME IS THE FLOORmissing precondition · required record absent · action out of scopevalue over threshold · stuck-detector fired · budget near exhaustiona verification step failed twicethe model is an added sensor — best for an instructionwith two readings that no rule could anticipateA ZERO RATE IS A WARNINGcount escalation separately fromfailure — or prompt authors and themodel both learn to avoid itzero means unreachable,not unnecessaryESCALATION IS A LEGITIMATE OUTCOME — TREAT IT AS SUCCESS IN THE METRICSand set an expiry on waiting escalations — a run parked indefinitely is a resource leak and a stale question nobody answers
Give the agent a schema'd escalation tool with runtime triggers as the floor, hand over context rather than a notification, and count escalation as success

Asking for Help Should Be an Available Action

An agent with no way to stop and ask will improvise, because continuing is the only action available to it. Give it an escalation tool with a schema — what it was trying to do, what it has established, what is blocking it, what it would do next if told to proceed, and what it needs from a person — and it becomes a legitimate outcome rather than a failure. This changes behaviour measurably in the right direction, but only if escalation is treated as success in your metrics. If the dashboard counts escalations against the agent, the pressure runs the wrong way, and both the prompt authors and the model will find ways to avoid them. Count escalation separately from failure, and watch the ratio: a rate of zero usually means the escape hatch is not reachable rather than that nothing has ever been ambiguous.

  • Without an escalation action, improvising is the only thing the agent can do
  • Give it a schema: goal, findings, blocker, proposed next action, what is needed
  • Count escalation separately from failure or the incentive pushes against using it
  • An escalation rate of zero usually means unreachable, not unnecessary

Where the Trigger Belongs

Escalation triggers should mostly live in the runtime rather than in the model's self-assessment, because a model asked to notice its own uncertainty is not a reliable instrument. The dependable triggers are structural: a precondition is missing, a required record does not exist, the action falls outside the permitted scope, a value exceeds a threshold, the stuck-detector fired, a budget is close to exhaustion, or a verification step failed twice. Model-initiated escalation is a valuable addition on top of these, particularly for genuine ambiguity in the request that no rule could anticipate — the instruction has two readings and they lead to different actions is exactly the case worth surfacing. Treat the runtime triggers as the floor and the model as an extra sensor, rather than relying on the model to know when it is out of its depth.

  • Structural triggers first: missing preconditions, scope, thresholds, stuck-detection, budget
  • A model assessing its own uncertainty is a weak instrument to rely on
  • Model-initiated escalation is genuinely valuable for ambiguity in the request itself
  • Runtime triggers are the floor; the model is an additional sensor on top

The Handover Determines Whether It Helps

An escalation that arrives as a notification saying the agent needs help has moved the work to a human without moving any of the context, and the person now has to reconstruct a run from a trace. Design the handover as a deliverable: the goal in the user's original words, what has been done so far and what state that leaves things in, the specific question, the options the agent sees with what it would choose and why, and a link to the full trace for anyone who wants it. Then make the response path direct — answering the question resumes the run from its checkpoint, rather than requiring someone to start a new one. Route by the decision required rather than to a general queue, and set an expectation for how long an escalation waits before it is abandoned, because a run parked indefinitely is a resource leak and a stale request nobody will answer.

  • A bare notification transfers the work without transferring the context
  • Hand over goal, progress, the specific question, the options, and the trace link
  • Answering should resume the checkpointed run, not require a fresh one
  • Route by decision type and set an expiry, or parked runs accumulate silently

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.