Beyond Vanilla Transformers: MoE and Friends
Mixture of Experts: Sparse by Design
A dense transformer activates every parameter for every token — expensive and increasingly unnecessary. Mixture-of-experts (MoE) replaces each feed-forward block with many parallel "expert" networks and a small router that sends each token to only a few of them. The result is a model with a very large total parameter count but a much smaller active parameter count per token. That decoupling is why most frontier models today are MoE: you get the knowledge capacity of a huge model at the inference compute of a much smaller one. It also explains why "how many parameters?" became a nearly meaningless question — total and active parameters can differ by an order of magnitude, and vendors rarely disclose either for proprietary models.
- Router: a learned gate picks the top-k experts per token, per layer
- Total parameters set capacity and memory footprint; active parameters set per-token compute
- Cost per token tracks active parameters — the economics behind cheap-but-capable frontier tiers
MoE Engineering Realities
MoE is not a free lunch. All experts must be resident in memory even though few run per token, so serving MoE demands large GPU memory or expert parallelism — sharding experts across devices and routing tokens over the interconnect. Training brings its own failure mode: without countermeasures the router collapses onto a few favourite experts while the rest go undertrained, so load-balancing losses and capacity limits are standard. Modern designs push further with many small experts, shared always-on experts, and fine-grained routing. For practitioners the takeaway is diagnostic: MoE explains why some models are cheap to run but expensive to host, and why batch throughput and latency behave differently than dense models of similar quality.
- Memory is provisioned for total parameters; compute is spent on active ones
- Load balancing prevents router collapse — an auxiliary training objective, not an afterthought
- Expert parallelism adds network traffic — interconnect bandwidth becomes a serving bottleneck
Cheaper Attention: Sliding Windows and Linear Variants
The second cost centre after the FFN is attention itself, and the same sparsity instinct applies: most tokens don't need to attend to everything. Sliding-window attention limits each token to a local neighbourhood, cutting cost from quadratic toward linear; information still travels far because each layer extends the effective receptive field. Many production models interleave local-attention layers with a few global ones — most layers cheap, a few layers with full reach. A separate research family (linear attention and friends) reformulates attention to avoid the all-pairs comparison entirely, trading some quality for constant-memory decoding. The pattern to internalise: full attention where it matters, cheaper approximations everywhere else — hybrids, not purism. These choices are invisible in the API but visible in pricing, speed, and long-context behaviour.
State-Space Models and Long-Context Engineering
The most radical alternative drops attention altogether. State-space models like Mamba process sequences recurrently through a fixed-size learned state: linear-time processing, constant memory during generation, no KV cache at all. Pure SSMs lag transformers on tasks needing precise recall of specific earlier tokens, so the practical direction is hybrids — mostly SSM or local-attention layers with a small number of full-attention layers for exact retrieval. Several serious open hybrid models have shipped, and the approach remains an active research direction rather than the frontier default. Long-context engineering ties this lesson together: RoPE scaling, sliding windows, GQA, cache compression, and hybrid layers are all facets of one question — how to make sequence length cheap without losing recall. When evaluating any long-context claim, watch effective context (measured recall across the window), not just the advertised length.
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.