The AI Learning Hub Journal

Multimodal Architecture

Multimodal understanding — the ingest pathImagepixelsVision encoderpatches → vectorsAudiowaveformAudio encoderframes → vectorsTextcharactersTokenizer + embedtokens → vectorsSharedrepresentationspaceone sequence ofvectors, whateverwent inTransformer modelreads every modality asone interleaved stream→ answer, caption, actionPATH 1 · BOLT-ON ADAPTERSFrozen text model + a trained projection layerCheap, fast to ship, easy to add a modalityModality is a guest: weaker cross-modal reasoningPATH 2 · NATIVELY MULTIMODALPretrained on interleaved image, audio and textExpensive: the data mix must be right up frontReasons across modalities, not merely about them
Every modality has to become vectors in one space — the open question is whether that space was designed in or bolted on

The Understanding Side: Encoders and Shared Space

For a language model to "see," images must become tokens the transformer can attend over. The dominant recipe: a vision encoder (typically a Vision Transformer) splits the image into patches, encodes each patch as a vector, and a projection layer maps those vectors into the same embedding space as text tokens. From the backbone's perspective, an image is just a run of unusual tokens interleaved with the text. Audio follows the same pattern — spectrogram frames or learned audio codecs become token sequences. The quality ceiling is set less by the encoder architecture than by the training data: models learn to connect what they see with what they read only when trained on genuinely interleaved image-text data at scale.

  • ViT patches → embeddings → projection into the language model's token space
  • Resolution matters: fixed low-resolution encoders miss small text and fine detail; tiling and native-resolution approaches address this
  • Audio: codec tokens or spectrogram features, same interleaving pattern
  • Attention is modality-blind — once everything is tokens, the backbone treats them uniformly

Native Multimodal vs Bolt-On

There are two ways to build a multimodal model, and the difference shows in behaviour. The bolt-on approach takes a trained text LLM, freezes or lightly tunes it, and grafts on a vision encoder with a projection layer — cheap, fast to ship, and how most early vision-language models were built. Native multimodal training instead feeds interleaved text, image, and audio data from early in pre-training, so cross-modal connections are learned deeply rather than adapted late. Native models tend to be markedly better at fine-grained visual reasoning — reading charts, interpreting documents, understanding spatial relationships — because vision was never a second-class input. Most current frontier models are natively multimodal on the understanding side; bolt-on remains common in the open-weight world because it lets a strong text model gain vision cheaply.

  • Bolt-on: pre-trained LLM + encoder + projection, aligned with a comparatively small training run
  • Native: interleaved multimodal data throughout pre-training — costlier, deeper integration
  • Symptom of bolt-on limits: fluent captions but brittle fine-grained reasoning (charts, dense documents, spatial layout)
  • Open-weight ecosystems favour bolt-on because it reuses existing strong text models

The Generation Side: Diffusion and Flow Matching

Generating images is a different problem from understanding them, and it grew up on a different architecture. Latent diffusion — the engine behind most image generators — works by compressing images into a smaller latent space with an autoencoder, then training a network to iteratively remove noise from random latents, guided by a text embedding of the prompt. Dozens of denoising steps gradually sculpt noise into a coherent image, which the decoder upsamples to pixels. Flow matching is the successor technique now standard in newer systems: instead of learning to reverse a noising process, the model learns a velocity field that transports noise to data along near-straight paths. In practice that means fewer sampling steps, more stable training, and faster generation at comparable quality — an engineering win more than a conceptual revolution.

  • Latent space: generate in a compressed representation, decode to pixels — orders of magnitude cheaper than pixel-space diffusion
  • Text conditioning: prompt embeddings steer every denoising step via cross-attention
  • Flow matching: learn straight-line noise-to-data transport; fewer steps, simpler training objective
  • Backbone shift: U-Nets have largely given way to diffusion transformers (DiT) that scale like LLMs

Autoregressive Generation and Video

A competing approach generates images the way LLMs generate text: quantise the image into discrete tokens and predict them one at a time with the same transformer that handles language. Autoregressive image generation has surged because it unifies understanding and generation in a single model — the model that reads your prompt, edits an image conversationally, and renders legible text inside the picture is exercising one set of weights, not calling out to a separate diffusion system. Video generation pushes the diffusion lineage hardest: diffusion transformers trained on video learn temporal consistency — objects persist, lighting stays coherent, physics roughly holds — which is why video models are increasingly discussed as nascent world models rather than mere clip generators. Compute cost remains the binding constraint; seconds of high-resolution video are still expensive to produce.

  • Autoregressive: image as token sequence — unified multimodal models generate and understand with shared weights
  • Unified models excel at instruction-following edits and text rendering; diffusion still competes on raw visual fidelity
  • Video diffusion transformers: temporal attention across frames enforces object and motion consistency
  • The world-model framing: predicting plausible video requires implicitly modelling how scenes evolve
Three generative familiesLatent diffusionStart: pure noisein a compressed latent spaceDenoise, many stepsU-Net or DiT, text-conditionedDecode latent → imagedecoder maps back to pixelsQuality scales with step count —each step is another model pass.Flow matchingStart: pure noisesame starting pointFollow a learned pathmodel predicts the velocityIntegrate → datastraighter path, fewer stepsSame destination as diffusion,reached in fewer, larger moves.AutoregressivePrompt → token gridimage or video as tokensPredict the next tokenone at a time, in orderDetokenize → framesdecoder rebuilds pixelsThe text-LLM recipe applied topixels; natural fit for video.All three learn a mapping from noise or tokens to data — the path taken is the design choice
Generation is a transport problem — the three families differ in the route from noise to data, not the destination

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.