Budgets and Cost Ceilings
Four Budgets, Not One
A step limit alone is insufficient, because the four things that can run away are only loosely correlated. Steps bound how many decisions the agent makes. Tokens bound context growth, and they grow superlinearly in a naive loop because every step re-sends everything before it. Wall-clock bounds how long a caller or a queue waits, and it is the one that matters for anything user-facing or anything holding a lock. Spend bounds the actual liability, and it is the only one denominated in the unit your finance team recognises. A run can sit comfortably inside a twenty-step limit and still consume a startling amount of money if each step carries an enormous context, or block a queue for an hour on a slow external tool. Track and enforce all four, per run, and treat any of them tripping as a distinct outcome.
- Steps, tokens, wall-clock and spend fail independently — bound each of them
- Token growth is superlinear in a naive loop: every step re-sends the history
- Wall-clock is the budget that matters for queues, locks and waiting callers
- Record which budget tripped; the four have completely different fixes
Enforce Before the Call, in the Runtime
Check budgets at the top of each iteration and before dispatching each tool call, not after the fact. Telling the model it has a limited number of steps in the system prompt is useful context and worthless enforcement — it will sometimes ration itself and sometimes not, and you cannot tell which. Budget state belongs in the run object, decremented by the runtime, and any subagent inherits from the parent's remaining allowance rather than receiving a fresh one, otherwise a delegating agent multiplies its own ceiling by the number of children it spawns. That specific bug is easy to write and expensive to discover. Per-tool limits are worth adding for anything with an external cost or a rate limit, so one expensive call cannot be issued forty times inside an otherwise reasonable run.
- Check at the top of the iteration and before each dispatch, never afterwards
- A budget mentioned in the prompt is context, not a control
- Subagents draw from the parent allowance — fresh budgets multiply the ceiling
- Add per-tool caps for anything externally expensive or rate limited
What Happens at the Ceiling
Hitting a limit is a designed outcome and should behave like one. The run should checkpoint its state, produce whatever partial result it legitimately has, record which budget tripped and at which step, and return a status the caller can act on rather than a generic failure. Silently truncating the context to keep going is the worst available option, because it converts a clean, attributable stop into an agent that continues with amnesia and produces confidently wrong work. Two refinements pay for themselves. Warn the model before the ceiling rather than at it — a step or two of notice lets it wrap up and write a useful handover instead of being cut mid-action. And make the resulting state resumable, so a human deciding the task deserves more budget can extend it rather than starting over.
- Checkpoint, return partial work, and report which budget tripped and where
- Silent truncation turns a clean stop into an agent working from amnesia
- Warn a step before the ceiling so the agent can write a handover
- Make the stopped state resumable — budget extension should not mean re-running
Setting the Numbers Honestly
Ceilings should come from the value and the tolerance of the task, not from whatever felt safe during development. Work out what a successful run is worth, decide what multiple of the median cost you are willing to pay before you would rather have a human look at it, and set the ceiling there. Then look at the distribution rather than the mean, because agent cost distributions have long tails and the tail is where both the money and the diagnostic value sit: runs at the ninety-fifth percentile of steps are usually confused rather than thorough, and reading a handful of them will tell you more about your tool descriptions than any aggregate metric. Expect to tighten ceilings over time as tools improve, and treat a rising median as a regression that needs an explanation even when the success rate is holding.
- Derive the ceiling from task value and the point where a human is cheaper
- Read the tail — high-step runs are confused far more often than they are thorough
- A rising median cost per success is a regression even at flat success rate
- Ceilings should tighten as tools and prompts improve, not drift upward
Try It Yourself
A ceiling you cannot point at in code is a ceiling you do not have, and the same is true of the terminal condition. This takes about half an hour on one agent and leaves you a table worth keeping in the design doc.
Pick one agent you own. If you own none, take an agent product you use regularly and fill this in for the design as you believe it works, or for the system you are currently sketching. Enumerate every way its loop can end: the terminal tools, the four budgets, and the stuck detector if you have one. For each row, write the number and the file and line where your code evaluates it. Then force one ceiling deliberately and check that the recorded reason is the row you predicted.
Agent: EXITS FROM THE LOOP Exit | Condition | Value | Enforced at (file:line) | Status returned Terminal tool - success | | | | Terminal tool - escalation | | | | Step budget | | | | Token budget | | | | Wall-clock budget | | | | Spend budget | | | | Stuck detector | | | | TERMINAL TOOL SCHEMA (required fields only) submit( result // the artefact the caller needs evidence // what was processed, what was checked, what the check returned identifiers // records touched caveats // what the caller must not assume ) Any row with no file and line is a condition you do not currently enforce.
- Every row names a file and a line, or is explicitly marked as unenforced
- You can point at the one line of code that ends the loop on success, and it is not a sentence in a prompt
- A deliberate overrun stops at the ceiling you predicted and reports which budget tripped
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.