The AI Learning Hub Journal

Cost and Latency as First-Class Metrics

Quality against cost and latency — and why the best is rarely the defaultaxes are qualitative on purpose — the shape holds even as the individual options changeQUALITYCOST AND LATENCY →the bar your task actually needssmall and fastmid-sizedstrongesteach further step buys lessWHY NOT SIMPLY ALWAYS USE THE BEST ONEMost requests are easya small model already answers them wellLatency is part of qualitya slow right answer can still miss the momentCost caps usageexpensive per call means fewer calls get madeQuality has a useful ceilingpast the bar your users need, the gains go unusedHard cases still existso the stronger option has to stay availableYou can have bothroute per request instead of choosing onceROUTING — THE DEFAULT BECOMES A DECISION MADE PER REQUESTIncomingrequestsRouterhow hard is this one?Easy → small, fast modelHard → stronger modelQuality where it countsat a fraction of the spendCachingidentical and near-identical asksshould not be paid for twiceBatchinggroup the work that is not urgentand run it when capacity is cheapTrim the contextevery unnecessary token is paid forin latency as well as in costDecide the quality bar first; everything after that is arithmetic about what clears it
The right default is the cheapest option that clears the bar — the strongest model is an escalation path

Measure Per Task, Not Per Call

Per-call token pricing is the wrong unit for reasoning about an AI system, because users do not buy calls — they complete tasks. The meaningful figure is total cost to complete one real task successfully, including retries, failed attempts, judge calls in the loop, and every step of an agent trajectory. Agents amplify the difference dramatically: a loop re-sends its growing context on every iteration, so a task that looks cheap per call can be expensive per completion, and a cheaper model that needs more steps or more retries can cost more overall than the expensive one. Report cost per successful task alongside quality in every eval run, and the model-selection conversation stops being about list prices and starts being about economics.

  • Cost per successful task, including retries and failures — not cost per call
  • Agent loops re-send context every step; per-call intuition breaks down
  • A cheaper model that needs more steps is not cheaper
  • Failed attempts are part of the cost of the successes

Latency Is Distribution, Not Average

Mean latency describes an experience nobody has. Report the distribution, and pay attention to the tail, because in a multi-step agent the tail dominates: a run with a dozen sequential steps inherits the slow tail of each one, and the probability that at least one step is slow rises quickly with step count. Time to first token is a different metric from time to completion and matters more for perceived responsiveness in interactive products, while for background automation only completion time matters and the whole optimisation changes. Latency is also an architectural constraint rather than an afterthought — it is what pushes you toward parallel tool calls, streaming progress, smaller models for routine steps, and speculative work that is discarded if unneeded.

  • Report p50, p95, p99 — tail latency is what users remember
  • Multi-step agents compound tail risk with every sequential step
  • Time to first token and time to completion are different products of different designs
  • Latency budgets drive architecture: parallelism, routing, streaming, smaller models

The Quality-Cost Frontier

Quality, cost, and latency trade against each other, and pretending otherwise leads to systems that are excellent and unaffordable. Evaluate configurations as points on a frontier rather than ranking them on quality alone: for each candidate — model, prompt strategy, number of reasoning steps, retrieval depth, judge usage — plot quality against cost per successful task and pick knowingly. The results are frequently unintuitive, because quality often flattens well before cost does, and a configuration a couple of points behind the best may cost a fraction as much. Routing exploits this directly: send easy inputs to a small model, escalate hard ones, and evaluate the routed system end to end. But evaluate the whole system, because a router that misclassifies difficulty can be worse and pricier than either model alone.

  • Plot quality against cost per successful task; choose a point deliberately
  • Quality usually flattens before cost does — the top configuration rarely wins on value
  • Cascades and routing exploit the frontier; evaluate the routed system end to end
  • Prompt-cache-friendly layout is a real cost lever in agent loops

Regressions in Cost Are Regressions

Teams gate on quality and let cost and latency drift, which is how a system quietly becomes twice as expensive over two quarters with no single change to blame. Put cost and latency in the same report as quality, with the same tiering and the same alerting, and treat a significant increase as a regression requiring justification even when quality improved. The usual culprits are cumulative and individually reasonable: a longer system prompt, an extra retrieved chunk, one more judge call, a retry policy that got more generous, a cache-invalidating token added early in the prompt. Each is defensible alone and the aggregate is not. Reviewing the cost diff on every change is a five-minute habit that prevents an expensive, hard-to-unwind surprise.

  • Report cost and latency next to quality on every run, with the same gates
  • Cost creep is cumulative and individually reasonable — review the diff per change
  • Watch for cache-invalidating edits early in prompts; they multiply across agent steps
  • A quality win that triples cost is a trade-off decision, not an automatic ship

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.