The Compute Stack
From Silicon to Model
Training and serving frontier models requires a coordinated stack — accelerators, memory, interconnect, networking, and distributed-training software — at a scale most engineers never touch directly. Clusters of tens of thousands to hundreds of thousands of accelerators must behave like one machine: a frontier training run is a single tightly-synchronised computation spread across a building. Understanding this stack explains model economics from first principles — why training runs cost what they do, why inference pricing keeps falling, why certain capabilities exist only behind cloud APIs, and why the industry's constraint has shifted from chip supply toward power, cooling, and datacentre buildout. It also explains the strategic behaviour of every major player: whoever controls the stack controls the margin.
- A frontier run is one synchronised computation across an entire datacentre
- Accelerators, memory, interconnect, and software each bottleneck at different scales
- Power and datacentre capacity have joined chips as first-order constraints
NVIDIA Blackwell: The Rack Is the Unit
NVIDIA's Blackwell generation — B200 chips, and GB200/GB300 systems pairing them with Grace CPUs — marks a shift in what the product actually is. The flagship offering is not a chip but a rack: GB200 NVL72 links 72 Blackwell GPUs over NVLink into what software can treat as one enormous accelerator with a unified fast-memory domain. That design exists because frontier models don't fit on any single device; the model is sharded across many GPUs, and the speed at which those shards exchange data governs everything. Blackwell's emphasis on low-precision formats (down to FP4 for inference) reflects where the money now is: serving models cheaply at volume, not just training them. NVIDIA's real moat is as much CUDA and its networking stack as the silicon itself — switching vendors means rewriting the software layer, which is why the ecosystem moves slowly even when alternatives look competitive on paper.
- GB200 NVL72: 72 GPUs in one NVLink domain — the rack as a single logical accelerator
- Low-precision inference formats (FP8/FP4) target serving economics, not just training speed
- CUDA plus networking lock-in, not raw FLOPS, sustains NVIDIA's position
TPUs and the Custom Silicon Trend
Google's TPU line is the longest-running proof that hyperscalers can escape the NVIDIA tax. TPU v6e (Trillium) targets efficient training and serving; v7 (Ironwood) is notable for being designed with inference as a first-class goal — a signal of where the market's compute demand has shifted. TPU pods connect thousands of chips via Google's own optical interconnect, and Google trains its Gemini family entirely on them. The same logic drives every large buyer toward custom silicon: Amazon's Trainium underpins Anthropic-scale training clusters, and Meta's MTIA serves internal inference workloads. The motivation is margin and supply security rather than beating NVIDIA outright — an accelerator that only needs to run your own workloads well can trade generality for cost. For builders the consequence is indirect but real: silicon diversity is a price lever, and serving stacks increasingly hide the hardware behind an API.
- TPU v6e Trillium and v7 Ironwood — Ironwood's inference-first design tracks the market's shift
- Trainium (Amazon) and MTIA (Meta): custom silicon for margin and supply security
- Vertical integration — own chip, own interconnect, own models — is the hyperscaler pattern
Why Memory Bandwidth and Interconnect Dominate
The counterintuitive truth of modern AI hardware: raw compute is rarely the bottleneck. LLM inference is memory-bound — generating each token requires streaming the model's weights and the KV cache through the chip, so tokens per second track memory bandwidth, not FLOPS. This is why every accelerator generation leads with HBM (high-bandwidth memory) capacity and bandwidth figures, and why HBM supply is one of the industry's tightest constraints. At cluster scale the same logic repeats one level up: training throughput depends on how fast thousands of chips exchange gradients and activations, making NVLink, InfiniBand-class fabrics, and optical interconnect the true differentiators. Read any hardware announcement with this lens — the memory and interconnect numbers tell you more than the compute numbers — and note that techniques you'll meet later (quantisation, KV-cache management, batching) are all attacks on the memory bottleneck.
- Token generation streams weights through memory every step — bandwidth sets the speed limit
- HBM capacity and supply are first-order constraints on the whole industry
- Cluster-scale interconnect determines whether ten thousand chips act like one machine
- Quantisation and KV caching are memory-bandwidth optimisations wearing software clothes
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.