The AI Learning Hub Journal
◆ Advanced

LLM Infrastructure

A trained model on disk does nothing — three layers get it to userseach layer solves a different problem, and each adds cost and complexity — read the stack from the bottom upthe userLAYER 3 — SERVING · RELIABLY TO USERS AT SCALEAPI gateway — rate limiting, authentication, routing, balancingrequest batching — concurrent requests share the GPUautoscaling — capacity follows demand, up and downmodel versioning — canary and blue-green rollouts, 5% of traffic firstthis is where enterprise SLAs liveLAYER 2 — OPTIMISATION · SMALLER, FASTER, CHEAPERquantisation — lower precision, less GPU memory, small quality lossdistillation — a compact model learns from a larger teacherpruning — removes low-impact weightsKV caching and speculative decoding — faster token generationthis layer sets the cost and latency in any commercial conversationLAYER 1 — HARDWARE · THE PHYSICAL COMPUTEGPUs and TPUs — NVIDIA and Google data-centre accelerator clustersdistributed training — a frontier model cannot fit on one machinecloud (AWS · Azure · GCP) rents the compute on demandon-premise — you buy the hardware: more control, far more costmost enterprises rent this layer; almost none own ittrained model on disk — does nothing by itselfWHEN THE COST QUESTION COMESlatency spikes usually happenhere, not inside the model —SLA requirements shape this layerask: full model or an optimisedvariant — and what qualitytradeoffs apply?top accelerators cost tens ofthousands each — a cluster is realcapital; quantify before on-premanswers for IT, procurement,and CISO infrastructure questionsHARDWARE RUNS IT, OPTIMISATION MAKES IT AFFORDABLE, SERVING KEEPS IT UPeach layer adds cost and complexity — the honest answer to "why is enterprise AI expensive?"
Three infrastructure layers stand between a trained model and its users — hardware runs it, optimisation makes it affordable, serving makes it reliable.

Three Layers Between the Model and the User

A trained model sitting on disk does nothing. Getting it to users reliably — at speed, at scale, and within cost — requires three distinct layers of infrastructure. Each layer solves a different problem, and each adds cost and complexity. Understanding this helps when customers ask why enterprise AI is expensive, or why on-premise deployment is harder than it sounds.

  • Hardware layer: the physical compute that trains and runs the model
  • Optimisation layer: techniques that make models smaller, faster, and cheaper to run
  • Serving layer: the infrastructure that gets model responses to users reliably at scale

Hardware — The Physical Compute

LLMs are computationally intense. The hardware needed to train and run them is specialised, expensive, and in short supply. Most enterprises never own this layer — they rent it through cloud providers. But knowing what it is helps when procurement or IT security asks about the underlying infrastructure.

  • GPUs and TPUs: the chips that power AI — NVIDIA H100 and A100 for general AI workloads, Google TPUv4 clusters for Google's own models
  • Distributed training: training splits across many nodes in parallel — a frontier model cannot fit on a single machine
  • Cloud vs on-premise: AWS, Azure, and GCP provide GPU compute on demand; on-premise means buying the hardware — higher control, much higher cost and complexity
  • H100s run ~$30K each; a small inference cluster is significant capital — use this when customers ask about on-prem AI cost

Optimisation — Making Models Smaller and Cheaper

Frontier models are too large and slow to run efficiently as-is. A set of techniques compresses and accelerates them without destroying their capability. This layer is what makes it practical to run powerful models at enterprise scale — and it directly affects the cost and latency numbers in any commercial conversation.

  • Quantisation: reduces model precision to save memory and cost — runs faster, uses less GPU memory with minimal quality loss
  • Distillation: a smaller model learns from a larger teacher model — compact model that behaves like the original at a fraction of the size
  • Pruning: removes low-impact weights to reduce model size without losing most capability
  • KV caching and speculative decoding: speeds up token generation at inference time — essential for conversational workloads

Serving — Getting the Model to Users Reliably

The serving layer is everything between the optimised model and the end user. It handles traffic, manages cost, and keeps the system running when demand spikes or something goes wrong. For enterprise deployments, this is where SLAs live.

  • API gateway: rate limiting, authentication, routing, and load balancing on every request
  • Request batching: groups concurrent requests for GPU efficiency — significantly cheaper than running them one at a time
  • Autoscaling: scales compute dynamically with demand — adds capacity when traffic spikes, reduces it when idle
  • Model versioning and canary deployments: safe rollout of model updates using canary and blue-green patterns — new versions go to a small percentage of traffic first

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.