The AI Learning Hub Journal
◆ Current Frontier

Reasoning Models

Standard model vs reasoning modelSTANDARD MODELPromptModelanswers immediatelyAnswerFast, cheap, flat costNo intermediate work to pay forREASONING MODELPromptINTERNAL THINKING TOKENS — billed, usually hiddenDecomposesplit the taskAttemptdraft a pathChecktest the workRevisefix and retryloops until the model is satisfied — length varies per questionAnswerWORTH THE COMPUTEMulti-step maths and proofsNon-trivial code and debuggingPlanning under many constraintsWRONG TOOLSimple, high-volume classificationLatency-sensitive interactive UXAnything a cheap model already tiesLatency and cost per call scale with how long it thinks — the same accuracy is a loss if the task was easy
Thinking is a spend, not a setting — buy it for hard problems, refuse it for cheap high-volume ones

From Prompting Trick to Trained Capability

Chain-of-thought started as a prompting technique — ask the model to "think step by step" and accuracy on hard problems improved. The breakthrough was making this a trained behaviour: use reinforcement learning on problems with verifiable answers (mathematics, code that either passes tests or doesn't) so the model learns which reasoning strategies actually lead to correct outcomes, not just which ones look plausible. RL-trained reasoning is now a standard capability class shipped by essentially every major vendor and by open-weight labs — it is table stakes, not a differentiator of any single model family. Most frontier models today are hybrids: the same model can answer immediately or deliberate first, with the amount of deliberation controllable per request.

  • Chain-of-thought prompting: elicits step-by-step reasoning but the model was never trained to reason well
  • RL on verifiable rewards: the model explores reasoning paths and is rewarded when the final answer checks out
  • Process supervision rewards good intermediate steps; outcome supervision rewards only the final answer — both are used in practice
  • The reasoning trace is a working scratchpad, not a faithful log — models can reach right answers via unfaithful traces

Test-Time Compute: The Second Scaling Axis

For years the only way to get a better model was a bigger training run. Reasoning models added a second axis: spend more compute at inference time on a specific problem. The same model produces meaningfully better answers on hard tasks when allowed to generate longer reasoning traces, explore alternatives, and check its own work. This reframes the engineering question from "which model?" to "how much thinking does this request deserve?" The tradeoff is direct: reasoning tokens cost money and add latency — often multiplying both by an order of magnitude — and the quality gains show diminishing returns. On easy tasks, extra thinking buys nothing at all; models can even "overthink" simple questions into worse answers.

  • More thinking helps most on maths, code, planning, and multi-step analysis — verifiable, decomposable problems
  • Thinking budgets: most APIs now let you cap or tune reasoning effort per request
  • Latency shifts from milliseconds to seconds or minutes — a product-design constraint, not just a cost line
  • Diminishing returns are real: the tenth thousand thinking tokens buy far less than the first thousand

When a Reasoning Model Is the Wrong Choice

Reasoning models are a tool, not an upgrade. A large share of production LLM traffic — classification, extraction, summarisation, formatting, routine chat — gains nothing from deliberation and pays the full latency and cost penalty for it. Interactive products with tight response-time budgets often cannot absorb multi-second thinking pauses at all. And in high-volume pipelines, a 10x token multiplier on every request is the difference between a viable unit economics and an unviable one. The practical pattern in production is routing: a fast model handles the bulk of traffic, and requests are escalated to extended reasoning only when the task is hard, the stakes are high, or the first attempt fails verification.

  • Simple, high-volume tasks: fast models match reasoning models at a fraction of cost and latency
  • Latency-sensitive UX (autocomplete, live chat, voice): thinking pauses break the product
  • Router architectures: classify request difficulty first, escalate selectively
  • Escalate on failure: try fast, verify, retry with reasoning — often cheaper than reasoning-first

Distilling Reasoning into Fast Models

The gap between "smart but slow" and "fast but shallow" is narrowing through distillation: generate reasoning traces and final answers with a large reasoning model, then train a smaller model on that output. The small model doesn't learn to deliberate at the same depth, but it absorbs much of the problem-solving behaviour — and runs at a fraction of the cost. Open-weight releases demonstrated this dramatically: compact distilled models now perform credibly on maths and code benchmarks that were frontier-only territory not long ago. For system designers, distillation is also a lifecycle strategy: prototype against the strongest model available, collect traces from your real workload, and distil a cheaper specialist for the narrow task you actually ship.

  • Teacher-student distillation: the large model's reasoning traces become the small model's training data
  • Distilled models inherit capability on the distribution they were trained on — they generalise less far off it
  • Workload-specific distillation: your production traces are a better curriculum than generic benchmarks
  • Hybrid deployments: distilled fast path plus frontier escalation covers most cost-quality frontiers

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.