Scaling Laws
The First Scaling Era: Bigger Pre-Training
Scaling laws are the empirical observation that language model loss falls as a smooth, predictable power law as you increase parameters, data, and compute. Kaplan et al. (2020) established the pattern and gave labs something the field had never had: a forecast. You could spend a small training run to predict what a run a thousand times larger would achieve. That predictability, not any single model, is what unlocked the capital of the current era — investors and labs could underwrite enormous training runs because the return on compute was no longer a guess. The first scaling era was defined by a simple recipe: scale pre-training and capabilities follow, including abilities that appear abruptly at scale thresholds rather than improving smoothly.
- Loss falls as a power law in parameters, data, and compute — smooth and forecastable
- Predictability is what made frontier-scale investment underwritable
- Some capabilities emerge at thresholds rather than improving gradually
- Primary source: "Scaling Laws for Neural Language Models" — the Kaplan paper that established the power-law fit
Chinchilla: The Data Correction
The Chinchilla result (Hoffmann et al., 2022) corrected the first era's bias toward parameters. For a fixed compute budget, most models of the time were substantially undertrained: compute-optimal training uses roughly 20 tokens per parameter, far more data per parameter than the field was using. A smaller model trained on more data beat larger, data-starved ones. The practical legacy is subtler than the headline: production models are now routinely trained well past the Chinchilla-optimal point, because Chinchilla optimises training compute alone. If you will serve a model billions of times, overtraining a smaller model is the better economic trade — inference cost scales with parameter count, and you pay it forever. Chinchilla-era parameter counts are public knowledge; note that frontier labs stopped disclosing parameter counts afterwards, so treat any figure you hear for current proprietary models as speculation.
- Chinchilla-optimal: roughly 20 training tokens per parameter for a fixed compute budget
- Production models overtrain past optimal because inference cost is paid on every request
- Frontier parameter counts are undisclosed — the era of public counts ended with Chinchilla-class models
- Primary source: "Training Compute-Optimal Large Language Models" — the Chinchilla paper
The Second Scaling Era: Inference-Time Compute
Around the arrival of reasoning models, a second scaling axis opened: spend more compute at inference time and get better answers. Instead of only making the model bigger before deployment, let it generate long internal reasoning traces, explore multiple solution paths, and check its own work before answering. Performance on hard problems — mathematics, code, multi-step planning — scales with thinking time in a way that echoes the original pre-training curves. This restructures the economics of capability: intelligence becomes a dial you turn per request rather than a fixed property of the weights. It also compounds with pre-training scaling rather than replacing it — a stronger base model gets more out of every thinking token. For builders, the immediate consequence is that cost and latency are now quality parameters you tune per task, not constants.
- Test-time compute: longer reasoning traces and search over solutions improve hard-task accuracy
- Capability becomes per-request and adjustable, not fixed at training time
- The two eras compound: better base models convert thinking tokens into capability more efficiently
The Data Wall
Pre-training scaling has an obvious dependency: an ever-larger supply of high-quality text. The stock of such data on the public internet is finite, and frontier training runs consume a meaningful fraction of it. This is the data wall argument — that data, not compute or capital, becomes the binding constraint on the first scaling era. The wall is best treated as an economic gradient rather than a cliff: each marginal token of quality data gets more expensive to find, license, or clean, which is why labs sign content deals, mine transcription and code, and invest heavily in data curation. Quality filtering turns out to matter as much as raw volume — aggressive deduplication and filtering can beat naive scale. The data wall is one of the main reasons the field pivoted toward the second scaling era: inference-time compute improves capability without requiring new human text.
- High-quality public text is finite; frontier runs already consume much of it
- Rising marginal cost of data, not sudden exhaustion, is the practical constraint
- Curation and deduplication substitute for volume — cleaner data trains better models
Synthetic Data and Distillation
Two responses to the data wall now sit at the centre of training pipelines. Synthetic data uses strong models to generate training material — reasoning traces, code with test-verified solutions, instruction-following examples — often filtered by automated checkers so only verified-correct samples survive. Done naively, training on your own outputs degrades quality; done with verification and diversity controls, it is how reasoning capability is now largely taught. Distillation transfers capability from a large teacher model into a smaller student, which is why compact open-weight models keep landing surprisingly close to frontier performance: they are compressing capability someone else paid to discover. Together these techniques mean capability now diffuses from frontier labs to the whole ecosystem faster than raw scaling alone would predict — a dynamic with obvious competitive and policy consequences.
- Verified synthetic data (checkable code, graded reasoning) sidesteps the quality-collapse risk
- Distillation compresses frontier capability into small, cheap-to-serve models
- Net effect: the capability gap between frontier and open-weight models closes faster than compute budgets suggest
Try It Yourself
The 20-tokens-per-parameter figure gets quoted as if it were a physical constant. Do the arithmetic on it, then check it against a model somebody actually trained, and you will see exactly what kind of claim it is.
Step one: for a 7-billion-parameter dense model, compute the compute-optimal training budget using the ~20 tokens-per-parameter heuristic, then estimate the training compute with the standard C ≈ 6ND approximation (6 FLOPs per parameter per token — roughly 2 for the forward pass and 4 for the backward pass). Step two: find an open-weight model whose card or technical report states both its parameter count and its training-token count, and compute its actual tokens-per-parameter ratio. Compare the two numbers and explain the gap using the inference-economics argument from the Chinchilla slide.
Heuristic budget: D = 20 × N N = 7,000,000,000 → D = [____] tokens Training compute: C ≈ 6 × N × D = [____] FLOPs Real model: N = [parameters] , D = [training tokens from its card] Actual ratio = D ÷ N = [____] tokens per parameter Simplifications: 6ND ignores attention FLOPs (which grow with context length) and counts dense parameters — for a mixture-of-experts model substitute active parameters per token, and note whether the card quotes total or active.
- D = 1.4 × 10^11 — 140 billion tokens — and C ≈ 5.9 × 10^21 FLOPs. If your exponent is off, you dropped a factor of 10 somewhere in the billions
- The model you looked up almost certainly has a ratio well above 20, often in the hundreds — that is deliberate overtraining, not a broken heuristic
- You can say why: 20:1 minimises training compute, and training compute is not what a model that will be served billions of times is optimised for
- Say it out loud once: this is an empirical fit from one family of training runs, not a law — it shifts with data quality, architecture, and which cost you are minimising
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.