The AI Learning Hub Journal

The Economics of Intelligence

The economics of intelligenceTotal spend over a model's lifecosttime →one-time training runlarge — but it happens onceinference spend accumulatesa little on every call, every day it is in servicetotal inference passes total trainingCost per token for the same capabilitycost / tokentime →same answer, steadily cheaperfalling unit cost is what makes the volume above possible
Training is a one-off spike; inference is a bill that never stops — even as each token gets cheaper

Training vs Inference: Two Different Businesses

AI economics splits cleanly in two. Training is capital expenditure: a huge, lumpy, one-time cost per model generation, paid by the vendor before any revenue, on hardware that depreciates fast. Inference is operating expenditure: paid on every request, forever, scaling directly with usage. The two have opposite improvement dynamics — training runs get more expensive each generation as frontier ambitions grow, while the cost of serving any fixed level of capability falls relentlessly. Every commercial structure in the industry falls out of this split: API pricing exists to amortise training capex across the widest possible usage base, model tiers exist because serving cost tracks model size, and vendors overtrain small models precisely because inference, not training, is where lifetime cost accumulates.

  • Training: lumpy vendor capex, paid before revenue, per model generation
  • Inference: metered opex, paid on every request, scales with adoption
  • Frontier training gets pricier per generation; serving fixed capability gets cheaper

Why Inference Now Dominates Spend

Early in the era, training budgets dominated the industry's ledger. As deployment scaled, aggregate inference spend overtook training — a model is trained once but queried billions of times, so success itself tilts the ratio. Two newer forces tilt it much further. Reasoning models convert inference compute into capability, multiplying tokens per query by large factors for hard tasks: thinking is metered. And agentic systems replace single question-answer exchanges with long tool-use loops — one user request can fan out into dozens of model calls. This is why the hardware market pivoted toward inference-optimised silicon, and why for builders cost engineering now means managing token consumption: caching, routing easy traffic to cheap models, bounding agent loops, and matching thinking budgets to task difficulty.

  • Deployment at scale flipped aggregate spend from training to inference
  • Reasoning tokens and agent loops multiply per-request consumption
  • Inference-first silicon (e.g. TPU Ironwood) is the hardware market confirming the shift

The Cost-per-Token Collapse

The price of any fixed level of capability has fallen by orders of magnitude within just a few years — one of the steepest sustained cost declines in computing history. The drivers stack: better hardware per dollar, low-precision serving, smarter batching and caching, distillation packing frontier-level capability into small models, and genuine price competition among providers, with open-weight releases acting as a hard ceiling on what anyone can charge. The strategic consequence for builders is that today's unit economics are a floor, not a fact: features that are marginally uneconomic at current prices become profitable on a short horizon. Teams that architect for this — routing by task difficulty, revisiting model choices quarterly — systematically outcompete teams that hard-code today's prices into their product decisions. The corollary discipline: falling unit costs invite exploding unit volumes, so spend governance still matters.

  • Orders-of-magnitude decline for fixed capability, driven by hardware, serving efficiency, and distillation
  • Open-weight availability caps API pricing power across the market
  • Design for tomorrow's prices: uneconomic features become viable on short horizons

Open Weights vs APIs — and What It Means for Builders

The API model sells convenience: zero infrastructure, frontier capability, per-token pricing, and someone else's ops team — at the cost of dependency on a provider's pricing, availability, and policies. Open-weight models (Llama, DeepSeek, Qwen and their successors) invert the trade: you pay for GPUs and expertise, gaining data locality, customisation, and unit costs that can beat API pricing at sustained high volume. The break-even is workload-shaped — spiky or low volume favours APIs; steady high-volume, latency-sensitive, or data-sovereign workloads favour self-hosting, with hosted open-weight providers as the middle path. In practice mature teams run portfolios: frontier APIs for the hardest tasks, small or self-hosted models for high-volume routine work. The durable builder posture is to hold your provider relationship loosely — abstract the model layer, benchmark alternatives regularly, and let the market's deflation work for you.

  • APIs: opex, zero infrastructure, provider dependency; open weights: capex plus ops, control and locality
  • Break-even depends on volume shape, latency needs, and data constraints — not ideology
  • Portfolio strategy plus a model-abstraction layer is the durable architecture

Try It Yourself

Almost every argument about whether an AI feature is worth building dissolves once someone computes the cost of one request and puts it next to the cost of the task it replaces. It is two lines of arithmetic and it is rarely done.

◆ Try it yourself

Pick one real task in your organisation that a model could do — triaging a ticket, drafting a first-pass summary, extracting fields from a document. Estimate its token shape: input tokens per request and output tokens per request. Look up current per-million input and output prices on your provider's pricing page (do not reuse a remembered number — these change often, almost always downward). Compute cost per request, then cost at your monthly volume, then the cost of a person doing the same task. The comparison, not either number alone, is the decision.

Cost per request = (in_tokens ÷ 1,000,000) × [price per 1M input]
                 + (out_tokens ÷ 1,000,000) × [price per 1M output]
  e.g. 3,000 in and 500 out  →  0.003 × [P_in] + 0.0005 × [P_out]

Honest cost   = cost per request × [retry factor, ≥ 1 for rejected or re-run outputs]
Monthly cost  = honest cost × [requests per month]

Human cost per task = [fully loaded hourly rate] × [minutes per task] ÷ 60
Minutes actually saved = [minutes per task] − [minutes of review that remain]
Break-even blended price per 1M = human cost per task ÷ (total tokens per request ÷ 1,000,000)
How you'll know it worked
  • 3,000 ÷ 1,000,000 = 0.003. If you wrote 0.03 or 3, your whole model is out by 10× or 1000× — this is the single most common slip
  • You priced input and output at their separate rates rather than one blended number
  • Your retry factor is above 1 and your saved-minutes figure subtracts the review that does not go away — a comparison without both is marketing, not analysis
  • The break-even price per million is probably far above what anyone charges today. That headroom, not the absolute cost, is the actual finding
  • Write the date next to your prices. Rerun this whenever you revisit the decision — the numbers move fast enough to change the answer

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.