The AI Learning Hub Journal

Cost and Latency as Requirements

Cost and latency are requirements, not reportsthey determine the architecture — a few seconds of tolerance cannot contain a ten-step sequential loopWRITE THE NUMBERS DOWN BEFORE YOU BUILDderive from the alternative: what the taskcosts when a person does it, and how longthe user or the queue will actually waitthen allocate: steps × context size × model tierTHE TAIL OF LONG RUNS DRIVES THE SPENDa small fraction of long, confused runsfrequently accounts for a large share ofthe bill — a step ceiling is a costcontrol, not only a safety oneWHERE THE MONEY GOES — EVERY STEP RE-SENDS THE ACCUMULATED CONTEXThistory, re-sent againnew this step123456input tokens dominate output by a widemargin — the naive loop's cost growswith roughly the square of run lengthCONTEXT DISCIPLINE OUTRANKS MODEL CHOICEas a cost lever — and the re-sent prefixis exactly the part caching can absorbattribute cost per step and per tool, or the one verbose tool result funding the bill stays invisibleLATENCY IS STRUCTURAL — SEQUENTIAL STEPS × TIME PER STEP; ONLY STRUCTURAL WINS ARE LARGEREMOVE STEPS FIRSTright-sized tools mean onecall does what three did —the biggest single winPARALLELISE IN A STEPindependent tool callsdispatched together, ontools safe to run at onceROUTE BY STEPa smaller model for themechanical steps; the largeone only for real decisionsMANAGE PERCEPTIONstream progress when a userwaits; honest status and anotification for backgroundMEASURE COST AND LATENCY PER SUCCESSFUL TASK — FAILURES AND RETRIES ARE PART OF THE PRICEderive the budget from the human alternative and real waiting tolerance, then allocate it across steps before building
Set cost per successful task and tail latency as budgets before building — input tokens dominate, so context discipline and step count are the levers

Write the Numbers Down Before You Build

Cost per successful task and latency at the tail are functional requirements for an agent, not reports you produce afterwards, because they determine the architecture rather than describing it. An interactive feature with a few seconds of tolerance cannot contain a ten-step sequential loop, and discovering that after building one means rebuilding. Derive the numbers from the alternative: what the task costs when a person does it, and how long the user or the queue will actually wait. Then decompose the budget across the loop — how many steps, what context size per step, which steps may use a larger model — so that the design conversation happens in terms of a budget you are allocating rather than a bill you receive. Track cost per successful task rather than per run, since failed and retried runs are part of what a success costs.

  • Cost per success and tail latency are requirements that determine the architecture
  • Derive them from the human alternative and from real waiting tolerance
  • Decompose the budget across steps, context size and model tier
  • Measure per successful task; failures and retries are part of the price of a success

Where the Money Actually Goes

Agent economics are dominated by input tokens rather than output, because every step re-sends the accumulated context, so a run's cost grows with roughly the square of its length in the naive case. Three consequences follow. Context discipline is the highest-leverage cost lever available, well ahead of model choice. Prompt caching matters disproportionately here, since the re-sent prefix is exactly the cacheable part. And the tail of the step distribution dominates the bill — a small fraction of long confused runs frequently accounts for a large share of spend, which is why a step ceiling is a cost control and not only a safety one. Instrument cost per step and per tool, not just per run, because that is what tells you whether one verbose tool result is quietly funding your entire cost problem.

  • Input tokens dominate: every step re-sends the history, so cost grows superlinearly
  • Context discipline outranks model choice as a cost lever
  • A small tail of long confused runs typically drives a large share of spend
  • Attribute cost per step and per tool, or the verbose tool result stays invisible

Latency Is Structural

Total latency is roughly the number of sequential steps multiplied by the time per step, which means the only large wins are structural. Reduce sequential steps by giving the model tools at the right granularity, so one call does what three did. Parallelise independent tool calls within a step, which requires tools designed to be safely concurrent and a loop that dispatches them together. Use a smaller model for mechanical steps where the judgement required is minimal, keeping the larger one for the decisions that actually need it. And manage perception where you cannot manage duration: streaming intermediate progress changes the experienced wait substantially for interactive features, and for background work, an honest status and a notification are usually worth more than shaving seconds. Optimise the step count first; per-call tuning rarely moves a multi-step run enough to matter.

  • Latency is sequential steps times time per step — attack the step count first
  • Right-sized tools remove steps; parallel dispatch removes waiting within a step
  • Route mechanical steps to a smaller model and keep the larger one for decisions
  • Stream progress for interactive work; status and notification for background work

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.