Quality Against Cost and Latency
Evaluate on the Frontier, Not on Quality Alone
Ranking configurations by quality alone answers a question nobody asked, because the best configuration is frequently one you cannot afford to run at volume. Evaluate on the frontier instead. For each candidate — model, prompt strategy, retrieval depth, number of reasoning steps, judge usage, retry policy — record quality, cost per successful task and latency at the tail together, then identify which options are not beaten on every axis at once. Everything else can be discarded without argument, and what remains is a small set of genuine trade-offs to choose between deliberately. The shape is usually the same and usually surprising: quality flattens well before cost does, so a configuration a little behind the leader can cost a fraction as much. Whether those points are worth the multiple is a product decision, and it can only be made when all three numbers appear on the same row of the same table.
- Record quality, cost per success and tail latency for every candidate together
- Discard options beaten on every axis; what remains is a real choice
- Quality typically flattens before cost does — the top option rarely wins on value
- All three numbers on one row, or the trade-off never actually gets discussed
Routing the Easy Cases Somewhere Cheaper
Most workloads are not uniformly hard. A large share of real traffic is routine and a small share carries most of the difficulty, which is what makes routing worth evaluating: send the easy cases to a smaller, faster model and escalate the rest. The mechanisms vary — a classifier over the input, a cheap first attempt with a validity or confidence check that triggers a retry on a stronger model, or explicit rules by case type — and they share one evaluation requirement. Measure the routed system end to end on the same suite, never the components in isolation, because a router that misclassifies difficulty produces answers worse than the small model would have given while costing more than the large one, having paid for both. Report the escalation rate as a metric in its own right and watch it over time, since drift in the traffic mix quietly changes the economics you signed off on.
- Traffic is not uniformly hard; routing exploits the difficulty distribution
- Classifier, cheap-attempt-with-check, or explicit rules — all need end-to-end evaluation
- A bad router costs more than either model alone and answers worse than both
- Track escalation rate over time; traffic drift changes the economics silently
Does the Expensive Option Earn Its Cost on Your Cases?
The general claim that a larger model is better is nearly always true and nearly always irrelevant, because the question is whether it is better on the cases you actually serve, by enough to justify the difference. Run the comparison properly: the same suite, prompts re-tuned per model rather than copied across, quality reported per slice instead of in aggregate, and cost measured per successful task including retries and failed attempts. The usual finding is that the advantage is concentrated — a large gap on a small hard slice and almost none across the bulk of traffic — which is an argument for routing rather than for standardising on one model everywhere. Then convert the difference into terms someone can decide with: what the gap costs per month at current volume, and what the failures it prevents would have cost. A quality difference nobody can price gets argued about indefinitely.
- The question is not which model is better, but better on your cases by how much
- Report per slice; the advantage is usually concentrated in one hard segment
- Cost per successful task, including retries and failures, is the comparable unit
- Price the gap at real volume — an unpriced difference is debated, never decided
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.