Inference & Latency Optimisation
The Shape of the Problem: Prefill vs Decode
LLM inference has two phases with opposite characters. Prefill processes the whole prompt in parallel — compute-bound, throughput-friendly, and responsible for time-to-first-token. Decode then generates one token at a time — each step must read the entire model's weights and the KV cache from GPU memory to produce a single token, making it memory-bandwidth-bound and painfully serial. This asymmetry explains most latency behaviour you observe: long prompts delay the first token, long outputs take time proportional to their length, and a GPU serving decode is mostly idle silicon waiting on memory. Every optimisation in this lesson attacks one side of this split: batch more work per memory read, make the memory footprint smaller, or cheat the serial nature of decoding itself.
- Time-to-first-token ≈ prefill cost; tokens-per-second ≈ decode cost — measure them separately
- Decode is memory-bound: the bottleneck is reading weights and KV cache, not arithmetic
- Streaming hides decode latency from users; it does not reduce it
Serving at Scale: Continuous Batching and PagedAttention
Naive serving batches requests together and waits for the longest one to finish — GPUs idle while short requests hold slots. Continuous batching schedules at the token level instead: every decode step, finished sequences leave the batch and queued requests join, keeping the GPU saturated. The second breakthrough was PagedAttention, introduced by vLLM: instead of reserving one contiguous memory region per request sized for the maximum possible context, the KV cache is split into small pages allocated on demand — virtual memory for attention. Fragmentation waste drops from the majority of KV memory to a few percent, which translates directly into larger batches and multiples of throughput. These two ideas are why vLLM and the serving engines that followed it became the default for self-hosted inference.
- Continuous batching: token-level scheduling — no request waits for a stranger to finish
- PagedAttention: paged, on-demand KV allocation — near-zero memory fragmentation
- Paged KV enables prefix sharing: identical prompt prefixes across requests reference the same pages
- Primary source: "Efficient Memory Management for Large Language Model Serving with PagedAttention" — the vLLM paper
Speculative Decoding: Cheating the Serial Bottleneck
Decode produces one token per expensive forward pass — unless you guess ahead. Speculative decoding uses a small, fast draft model to propose several tokens, then the large model verifies the whole draft in a single parallel pass. Accepted tokens are guaranteed to match exactly what the large model would have produced alone — the math preserves the output distribution, so there is no quality cost, only a speed gain that depends on how often the draft is right. Predictable text (code, boilerplate, structured output) accepts long drafts and speeds up dramatically; surprising text accepts little. Variants replace the separate draft model with extra prediction heads on the main model (Medusa-style) or n-gram lookup. Most large providers run some form of this silently — it is part of why API speeds keep improving on unchanged models.
Quantisation: GPTQ, AWQ, GGUF
Weights trained in 16-bit precision carry more precision than inference needs. Quantisation stores them in 8 or 4 bits, shrinking memory footprint several-fold — and since decode is memory-bound, smaller weights also mean faster tokens. The mainstream methods are post-training: GPTQ quantises layer by layer while correcting the error each layer introduces; AWQ observes activations to find the small fraction of weights that matter most and protects them, quantising the rest aggressively. GGUF is not a method but a file format — the llama.cpp ecosystem's container with a menu of quantisation levels for CPU and consumer-GPU inference, and the reason capable models run on laptops at all. Quality loss at 8-bit is negligible and at careful 4-bit usually small, but it is task-dependent — quantised models must be evaluated on your workload, not assumed equivalent.
- Memory-bound decode means quantisation buys speed as well as capacity
- GPTQ: error-compensating layer-wise quantisation; AWQ: protect activation-critical weights; GGUF: the packaging format behind local inference
- Evaluate quantised models on your own tasks — degradation concentrates unpredictably (often math and edge-case reasoning)
Try It Yourself
Serving capacity is usually decided by KV cache memory, not by the model weights — and the arithmetic is small enough to do on paper. Work it once and you will never again be surprised by a concurrency limit.
Compute KV cache size for a hypothetical model: 32 layers, 8 key/value heads, head dimension 128, cache held in FP16 (2 bytes per element), at a context length of 8,192 tokens. Get bytes per token first, then per sequence, then per server at 50 concurrent sequences. Then redo it with the configuration of a model you would actually deploy — an open-weight model card gives you layers, head count and head dimension. Divide bytes by 1,073,741,824 for GiB.
Per token = 2 × [layers] × [kv_heads] × [head_dim] × [bytes_per_element] Per sequence = per-token × [context_tokens] Per server = per-sequence × [concurrent_sequences] Headroom check: [accelerator memory] − [model weights] = memory left for KV Max concurrency ≈ memory left ÷ per-sequence Simplifications: the 2 counts K and V; assumes every layer is full attention with the same KV-head count; ignores allocator and paging overhead, sliding-window layers, and any KV compression or cache quantisation.
- Per token you should get 131,072 bytes — 128 KiB — and per sequence at 8,192 tokens exactly 1 GiB
- 50 concurrent sequences is 50 GiB of KV alone; subtract the weights from your accelerator memory and see how little headroom is left
- Set kv_heads to 32 instead of 8 (the no-GQA case): the cache is 4× larger, which is the whole argument for grouped-query attention in one number
- Doubling context doubles the cache and no more — KV grows linearly with length; it is prefill compute, not cache memory, that grows quadratically
- Halving the element size (FP16 to FP8) halves the cache — the same lever as weight quantisation, applied to the other memory consumer
Turn sampling up and down and watch the output change.
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.