The AI Learning Hub Journal

The Compute Stack

The AI compute stack — and where it actually stallsAcceleratorsGPU · TPU · custom siliconpeak FLOPs, tensor cores, low-precision formatsMemoryHBM capacity + bandwidthweights and KV cache must be fed on every stepBOTTLENECKInterconnectintra-node links · cluster fabriccollective ops set the pace of a distributed runBOTTLENECKCluster & datacenterpower · cooling · floor spacecapacity is gated by megawatts, not by chip countWhy FLOPs misleadPeak FLOPs is a spec-sheet number.Real throughput is set by how fastweights reach the compute units.Decode is bandwidth-bound: everytoken re-reads the whole model.Scale-out is link-bound: the slowesthop throttles the entire cluster.Rule of thumbA busy accelerator is oftenan idle one, waiting on data.Compare bandwidth and topology,not headline FLOPs.Every layer must keep the one above it fed — the stack runs at the speed of its narrowest pipe.Memory bandwidth per dollar and interconnect topology decide real-world performance.
Accelerators rarely run out of math first — they run out of bytes delivered per second

From Silicon to Model

Training and serving frontier models requires a coordinated stack — accelerators, memory, interconnect, networking, and distributed-training software — at a scale most engineers never touch directly. Clusters of tens of thousands to hundreds of thousands of accelerators must behave like one machine: a frontier training run is a single tightly-synchronised computation spread across a building. Understanding this stack explains model economics from first principles — why training runs cost what they do, why inference pricing keeps falling, why certain capabilities exist only behind cloud APIs, and why the industry's constraint has shifted from chip supply toward power, cooling, and datacentre buildout. It also explains the strategic behaviour of every major player: whoever controls the stack controls the margin.

  • A frontier run is one synchronised computation across an entire datacentre
  • Accelerators, memory, interconnect, and software each bottleneck at different scales
  • Power and datacentre capacity have joined chips as first-order constraints

NVIDIA Blackwell: The Rack Is the Unit

NVIDIA's Blackwell generation — B200 chips, and GB200/GB300 systems pairing them with Grace CPUs — marks a shift in what the product actually is. The flagship offering is not a chip but a rack: GB200 NVL72 links 72 Blackwell GPUs over NVLink into what software can treat as one enormous accelerator with a unified fast-memory domain. That design exists because frontier models don't fit on any single device; the model is sharded across many GPUs, and the speed at which those shards exchange data governs everything. Blackwell's emphasis on low-precision formats (down to FP4 for inference) reflects where the money now is: serving models cheaply at volume, not just training them. NVIDIA's real moat is as much CUDA and its networking stack as the silicon itself — switching vendors means rewriting the software layer, which is why the ecosystem moves slowly even when alternatives look competitive on paper.

  • GB200 NVL72: 72 GPUs in one NVLink domain — the rack as a single logical accelerator
  • Low-precision inference formats (FP8/FP4) target serving economics, not just training speed
  • CUDA plus networking lock-in, not raw FLOPS, sustains NVIDIA's position

TPUs and the Custom Silicon Trend

Google's TPU line is the longest-running proof that hyperscalers can escape the NVIDIA tax. TPU v6e (Trillium) targets efficient training and serving; v7 (Ironwood) is notable for being designed with inference as a first-class goal — a signal of where the market's compute demand has shifted. TPU pods connect thousands of chips via Google's own optical interconnect, and Google trains its Gemini family entirely on them. The same logic drives every large buyer toward custom silicon: Amazon's Trainium underpins Anthropic-scale training clusters, and Meta's MTIA serves internal inference workloads. The motivation is margin and supply security rather than beating NVIDIA outright — an accelerator that only needs to run your own workloads well can trade generality for cost. For builders the consequence is indirect but real: silicon diversity is a price lever, and serving stacks increasingly hide the hardware behind an API.

  • TPU v6e Trillium and v7 Ironwood — Ironwood's inference-first design tracks the market's shift
  • Trainium (Amazon) and MTIA (Meta): custom silicon for margin and supply security
  • Vertical integration — own chip, own interconnect, own models — is the hyperscaler pattern

Why Memory Bandwidth and Interconnect Dominate

The counterintuitive truth of modern AI hardware: raw compute is rarely the bottleneck. LLM inference is memory-bound — generating each token requires streaming the model's weights and the KV cache through the chip, so tokens per second track memory bandwidth, not FLOPS. This is why every accelerator generation leads with HBM (high-bandwidth memory) capacity and bandwidth figures, and why HBM supply is one of the industry's tightest constraints. At cluster scale the same logic repeats one level up: training throughput depends on how fast thousands of chips exchange gradients and activations, making NVLink, InfiniBand-class fabrics, and optical interconnect the true differentiators. Read any hardware announcement with this lens — the memory and interconnect numbers tell you more than the compute numbers — and note that techniques you'll meet later (quantisation, KV-cache management, batching) are all attacks on the memory bottleneck.

  • Token generation streams weights through memory every step — bandwidth sets the speed limit
  • HBM capacity and supply are first-order constraints on the whole industry
  • Cluster-scale interconnect determines whether ten thousand chips act like one machine
  • Quantisation and KV caching are memory-bandwidth optimisations wearing software clothes

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.