Cost and Latency as First-Class Metrics
Measure Per Task, Not Per Call
Per-call token pricing is the wrong unit for reasoning about an AI system, because users do not buy calls — they complete tasks. The meaningful figure is total cost to complete one real task successfully, including retries, failed attempts, judge calls in the loop, and every step of an agent trajectory. Agents amplify the difference dramatically: a loop re-sends its growing context on every iteration, so a task that looks cheap per call can be expensive per completion, and a cheaper model that needs more steps or more retries can cost more overall than the expensive one. Report cost per successful task alongside quality in every eval run, and the model-selection conversation stops being about list prices and starts being about economics.
- Cost per successful task, including retries and failures — not cost per call
- Agent loops re-send context every step; per-call intuition breaks down
- A cheaper model that needs more steps is not cheaper
- Failed attempts are part of the cost of the successes
Latency Is Distribution, Not Average
Mean latency describes an experience nobody has. Report the distribution, and pay attention to the tail, because in a multi-step agent the tail dominates: a run with a dozen sequential steps inherits the slow tail of each one, and the probability that at least one step is slow rises quickly with step count. Time to first token is a different metric from time to completion and matters more for perceived responsiveness in interactive products, while for background automation only completion time matters and the whole optimisation changes. Latency is also an architectural constraint rather than an afterthought — it is what pushes you toward parallel tool calls, streaming progress, smaller models for routine steps, and speculative work that is discarded if unneeded.
- Report p50, p95, p99 — tail latency is what users remember
- Multi-step agents compound tail risk with every sequential step
- Time to first token and time to completion are different products of different designs
- Latency budgets drive architecture: parallelism, routing, streaming, smaller models
The Quality-Cost Frontier
Quality, cost, and latency trade against each other, and pretending otherwise leads to systems that are excellent and unaffordable. Evaluate configurations as points on a frontier rather than ranking them on quality alone: for each candidate — model, prompt strategy, number of reasoning steps, retrieval depth, judge usage — plot quality against cost per successful task and pick knowingly. The results are frequently unintuitive, because quality often flattens well before cost does, and a configuration a couple of points behind the best may cost a fraction as much. Routing exploits this directly: send easy inputs to a small model, escalate hard ones, and evaluate the routed system end to end. But evaluate the whole system, because a router that misclassifies difficulty can be worse and pricier than either model alone.
- Plot quality against cost per successful task; choose a point deliberately
- Quality usually flattens before cost does — the top configuration rarely wins on value
- Cascades and routing exploit the frontier; evaluate the routed system end to end
- Prompt-cache-friendly layout is a real cost lever in agent loops
Regressions in Cost Are Regressions
Teams gate on quality and let cost and latency drift, which is how a system quietly becomes twice as expensive over two quarters with no single change to blame. Put cost and latency in the same report as quality, with the same tiering and the same alerting, and treat a significant increase as a regression requiring justification even when quality improved. The usual culprits are cumulative and individually reasonable: a longer system prompt, an extra retrieved chunk, one more judge call, a retry policy that got more generous, a cache-invalidating token added early in the prompt. Each is defensible alone and the aggregate is not. Reviewing the cost diff on every change is a five-minute habit that prevents an expensive, hard-to-unwind surprise.
- Report cost and latency next to quality on every run, with the same gates
- Cost creep is cumulative and individually reasonable — review the diff per change
- Watch for cache-invalidating edits early in prompts; they multiply across agent steps
- A quality win that triples cost is a trade-off decision, not an automatic ship
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.