The AI Learning Hub Journal
◆ Architecture

The Transformer Architecture

phishing emailcredential harvestingBEC fraudmalware sampleransomware payloadtrojan dropperfirewall lognetwork telemetryflow dataIdentity attack clusterMalware clusterNetwork telemetry cluster2D projection of high-dimensional embedding space
Semantically similar concepts cluster together — embeddings power similarity search

Attention Is All You Need

The transformer, introduced by Vaswani et al. in 2017, replaced recurrence with self-attention — every token can attend to every other token in the sequence in a single parallel operation. That one change removed the sequential bottleneck of RNNs and made it economical to train on essentially unlimited text with massive GPU clusters. Nearly a decade later, every frontier language model is still a transformer at its core, though heavily modified: different attention variants, different position encodings, different feed-forward blocks. The skeleton is unchanged: embed tokens, alternate attention and feed-forward layers with residual connections, project back to vocabulary logits. Understanding this skeleton is the prerequisite for everything else in this module — inference cost, context limits, and fine-tuning all trace back to it.

  • Each block: attention (tokens exchange information) then a feed-forward network (each token processed independently)
  • The final layer projects to a probability distribution over the vocabulary — one next token at a time

Attention Mechanics: Queries, Keys, Values

Attention is a soft lookup. Each token produces three vectors: a query ("what am I looking for?"), a key ("what do I contain?"), and a value ("what do I contribute?"). Every query is compared against every key via dot product, the scores are scaled and softmaxed into weights, and the output is the weighted sum of values. Multi-head attention runs this many times in parallel with different learned projections, letting different heads specialise — some track syntax, some track coreference, some copy verbatim. The cost is the catch: comparing all pairs is quadratic in sequence length, which is why raw attention over very long inputs is expensive and why so much engineering effort targets exactly this bottleneck.

  • Attention(Q, K, V) = softmax(QKᵀ/√d)V — one line of math powering the whole field
  • Multi-head: parallel attention with separate projections, concatenated and mixed
  • Quadratic scaling: doubling context length quadruples attention compute in prefill

Position: RoPE and the Road to Long Context

Attention is order-agnostic — "dog bites man" and "man bites dog" would be identical without position information. Early transformers added absolute position vectors to each embedding, but that ties the model to the lengths it saw in training. Modern LLMs almost universally use RoPE (rotary position embeddings): positions are encoded by rotating query and key vectors by an angle proportional to their index, so attention naturally depends on relative distance between tokens. RoPE also extends: by rescaling the rotation frequencies (position interpolation, YaRN and successors) plus targeted long-context training, models trained mostly on shorter sequences can operate at hundreds of thousands of tokens. That is the main reason context windows grew so quickly — it was a position-encoding and data problem as much as a raw compute problem. Even so, effective use of long context lags claimed window size: retrieval quality inside the window degrades ("lost in the middle").

The KV Cache: Why Generation Works the Way It Does

Generation is autoregressive: the model produces one token, appends it, and runs again. Naively, every step would recompute attention over the entire sequence. The KV cache fixes this — keys and values for all previous tokens are computed once and stored, so each new token only computes its own query against the cached history. This makes decoding fast but moves the cost into GPU memory: the cache grows linearly with context length and can rival the model weights for long conversations. That memory pressure drove architectural responses — multi-query and grouped-query attention (GQA) share key/value heads across query heads, shrinking the cache several-fold with minimal quality loss, and most modern open and frontier models use some variant. When a provider offers prompt caching, it is literally persisting this KV cache between requests.

  • Prefill (process the prompt, parallel) and decode (one token at a time, sequential) are distinct phases with different bottlenecks
  • KV cache trades memory for compute — long contexts are memory-hungry, not just slow
  • GQA: many query heads share fewer KV heads — a near-free cache reduction
◆ See it for yourself
Open Next-Word Prediction in the library →

Watch the model choose the next word, one at a time.

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.