The AI Learning Hub Journal

Beyond Vanilla Transformers: MoE and Friends

Mixture of experts — sparse activationtoken 1token 2token 3tokens inRoutergating networkpicks top-k of Nper token, per layerExpert 1 — idle, residentExpert 2 — idle, residentExpert 3 — activeExpert 4 — idle, residentExpert 5 — idle, residentExpert 6 — activeExpert 7 — idle, residentExpert 8 — idle, residentN experts, all held in memoryWeighted sum→ next layerParameters resident in memory vs parameters used for one tokenactive per tokentotal parameters — every expert stays loaded in HBM
Memory cost scales with every expert; compute cost scales only with the few that light up

Mixture of Experts: Sparse by Design

A dense transformer activates every parameter for every token — expensive and increasingly unnecessary. Mixture-of-experts (MoE) replaces each feed-forward block with many parallel "expert" networks and a small router that sends each token to only a few of them. The result is a model with a very large total parameter count but a much smaller active parameter count per token. That decoupling is why most frontier models today are MoE: you get the knowledge capacity of a huge model at the inference compute of a much smaller one. It also explains why "how many parameters?" became a nearly meaningless question — total and active parameters can differ by an order of magnitude, and vendors rarely disclose either for proprietary models.

  • Router: a learned gate picks the top-k experts per token, per layer
  • Total parameters set capacity and memory footprint; active parameters set per-token compute
  • Cost per token tracks active parameters — the economics behind cheap-but-capable frontier tiers

MoE Engineering Realities

MoE is not a free lunch. All experts must be resident in memory even though few run per token, so serving MoE demands large GPU memory or expert parallelism — sharding experts across devices and routing tokens over the interconnect. Training brings its own failure mode: without countermeasures the router collapses onto a few favourite experts while the rest go undertrained, so load-balancing losses and capacity limits are standard. Modern designs push further with many small experts, shared always-on experts, and fine-grained routing. For practitioners the takeaway is diagnostic: MoE explains why some models are cheap to run but expensive to host, and why batch throughput and latency behave differently than dense models of similar quality.

  • Memory is provisioned for total parameters; compute is spent on active ones
  • Load balancing prevents router collapse — an auxiliary training objective, not an afterthought
  • Expert parallelism adds network traffic — interconnect bandwidth becomes a serving bottleneck

Cheaper Attention: Sliding Windows and Linear Variants

The second cost centre after the FFN is attention itself, and the same sparsity instinct applies: most tokens don't need to attend to everything. Sliding-window attention limits each token to a local neighbourhood, cutting cost from quadratic toward linear; information still travels far because each layer extends the effective receptive field. Many production models interleave local-attention layers with a few global ones — most layers cheap, a few layers with full reach. A separate research family (linear attention and friends) reformulates attention to avoid the all-pairs comparison entirely, trading some quality for constant-memory decoding. The pattern to internalise: full attention where it matters, cheaper approximations everywhere else — hybrids, not purism. These choices are invisible in the API but visible in pricing, speed, and long-context behaviour.

State-Space Models and Long-Context Engineering

The most radical alternative drops attention altogether. State-space models like Mamba process sequences recurrently through a fixed-size learned state: linear-time processing, constant memory during generation, no KV cache at all. Pure SSMs lag transformers on tasks needing precise recall of specific earlier tokens, so the practical direction is hybrids — mostly SSM or local-attention layers with a small number of full-attention layers for exact retrieval. Several serious open hybrid models have shipped, and the approach remains an active research direction rather than the frontier default. Long-context engineering ties this lesson together: RoPE scaling, sliding windows, GQA, cache compression, and hybrid layers are all facets of one question — how to make sequence length cheap without losing recall. When evaluating any long-context claim, watch effective context (measured recall across the window), not just the advertised length.

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.