Multimodal Architecture
The Understanding Side: Encoders and Shared Space
For a language model to "see," images must become tokens the transformer can attend over. The dominant recipe: a vision encoder (typically a Vision Transformer) splits the image into patches, encodes each patch as a vector, and a projection layer maps those vectors into the same embedding space as text tokens. From the backbone's perspective, an image is just a run of unusual tokens interleaved with the text. Audio follows the same pattern — spectrogram frames or learned audio codecs become token sequences. The quality ceiling is set less by the encoder architecture than by the training data: models learn to connect what they see with what they read only when trained on genuinely interleaved image-text data at scale.
- ViT patches → embeddings → projection into the language model's token space
- Resolution matters: fixed low-resolution encoders miss small text and fine detail; tiling and native-resolution approaches address this
- Audio: codec tokens or spectrogram features, same interleaving pattern
- Attention is modality-blind — once everything is tokens, the backbone treats them uniformly
Native Multimodal vs Bolt-On
There are two ways to build a multimodal model, and the difference shows in behaviour. The bolt-on approach takes a trained text LLM, freezes or lightly tunes it, and grafts on a vision encoder with a projection layer — cheap, fast to ship, and how most early vision-language models were built. Native multimodal training instead feeds interleaved text, image, and audio data from early in pre-training, so cross-modal connections are learned deeply rather than adapted late. Native models tend to be markedly better at fine-grained visual reasoning — reading charts, interpreting documents, understanding spatial relationships — because vision was never a second-class input. Most current frontier models are natively multimodal on the understanding side; bolt-on remains common in the open-weight world because it lets a strong text model gain vision cheaply.
- Bolt-on: pre-trained LLM + encoder + projection, aligned with a comparatively small training run
- Native: interleaved multimodal data throughout pre-training — costlier, deeper integration
- Symptom of bolt-on limits: fluent captions but brittle fine-grained reasoning (charts, dense documents, spatial layout)
- Open-weight ecosystems favour bolt-on because it reuses existing strong text models
The Generation Side: Diffusion and Flow Matching
Generating images is a different problem from understanding them, and it grew up on a different architecture. Latent diffusion — the engine behind most image generators — works by compressing images into a smaller latent space with an autoencoder, then training a network to iteratively remove noise from random latents, guided by a text embedding of the prompt. Dozens of denoising steps gradually sculpt noise into a coherent image, which the decoder upsamples to pixels. Flow matching is the successor technique now standard in newer systems: instead of learning to reverse a noising process, the model learns a velocity field that transports noise to data along near-straight paths. In practice that means fewer sampling steps, more stable training, and faster generation at comparable quality — an engineering win more than a conceptual revolution.
- Latent space: generate in a compressed representation, decode to pixels — orders of magnitude cheaper than pixel-space diffusion
- Text conditioning: prompt embeddings steer every denoising step via cross-attention
- Flow matching: learn straight-line noise-to-data transport; fewer steps, simpler training objective
- Backbone shift: U-Nets have largely given way to diffusion transformers (DiT) that scale like LLMs
Autoregressive Generation and Video
A competing approach generates images the way LLMs generate text: quantise the image into discrete tokens and predict them one at a time with the same transformer that handles language. Autoregressive image generation has surged because it unifies understanding and generation in a single model — the model that reads your prompt, edits an image conversationally, and renders legible text inside the picture is exercising one set of weights, not calling out to a separate diffusion system. Video generation pushes the diffusion lineage hardest: diffusion transformers trained on video learn temporal consistency — objects persist, lighting stays coherent, physics roughly holds — which is why video models are increasingly discussed as nascent world models rather than mere clip generators. Compute cost remains the binding constraint; seconds of high-resolution video are still expensive to produce.
- Autoregressive: image as token sequence — unified multimodal models generate and understand with shared weights
- Unified models excel at instruction-following edits and text rendering; diffusion still competes on raw visual fidelity
- Video diffusion transformers: temporal attention across frames enforces object and motion consistency
- The world-model framing: predicting plausible video requires implicitly modelling how scenes evolve
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.