Structured Output and Constrained Decoding
Constraining the Shape of What the Model Can Emit
Constrained decoding restricts generation so the output must conform to a schema or grammar: at each step, tokens that would make the result invalid are excluded, so the model cannot produce malformed structure even if it would otherwise. As a security control this matters because it removes an entire class of downstream failure. When output is guaranteed to be a valid object with typed fields, the consumer does not need to parse free text, and an injected instruction cannot introduce structure the consumer will misread — no extra fields, no trailing content, no format confusion. It also shrinks the expressive room available to a persuaded model: if a tool takes an enumerated action and a validated identifier, the space of harmful calls is bounded by the schema regardless of what the model was persuaded to attempt. Constrain the shape, and you constrain the reachable outcomes.
- Generation is restricted to tokens that keep the output schema-valid
- Removes format confusion and parsing ambiguity at the consumer
- Enumerated actions and validated identifiers bound what a persuaded model can attempt
- Constraining shape constrains the set of reachable outcomes
The Limit: Valid Is Not Safe
The essential caveat is that a schema constrains form, not intent. A transfer object with a valid recipient identifier and a valid amount is schema-perfect and may still be the wrong recipient and the wrong amount, chosen because a document told the model to choose them. Every free-text field inside a valid object is an unconstrained channel — a message body, a note, a filename, a search query — and if any consumer renders or executes that field, the constraint has bought nothing there. Schemas are also frequently permissive in practice: an identifier typed as a string rather than a pattern-validated reference, an enum that includes an unused destructive option, an object with an open additional-properties policy. Constrained decoding is a strong hygiene control that eliminates a class of parsing failures. It is not authorisation, and treating it as one is a common overreach.
- Schemas constrain form; the values inside remain attacker-influenceable
- Free-text fields inside a valid object are unconstrained channels
- Permissive schemas — loose string types, unused destructive enum values — give back the ground
- Structure is not authorisation; the policy check still has to happen
Validate at Every Consumer
The complementary control is proper output handling wherever model output is used, and it follows ordinary application security practice applied to a source many teams forget to treat as untrusted. Text rendered in a browser needs encoding and a policy that prevents automatic fetching of remote resources referenced in it. Content passed to a shell, a query engine, a template renderer, or a deserialiser needs the same treatment as any user-supplied input, because that is exactly what it is. Identifiers should be resolved and authorised at the consumer rather than trusted because they arrived in a typed field. And output consumed by another agent should be treated as data unless you have deliberately decided otherwise. The recurring failure is a pipeline where output was validated once at the model boundary and then treated as internal by five components downstream.
- Model output is untrusted input at every consumer, not only at the first one
- Encode for the rendering context and disable automatic remote resource fetching
- Resolve and authorise identifiers at the consumer, not on arrival
- Validate once at the boundary and trust everywhere after is the pattern to avoid
- OWASP's LLM Top 10 names this downstream failure improper output handling — treat every consumer as the enforcement point
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.