The AI Learning Hub Journal

Structured Output and Constrained Decoding

Constrain the shape, and you constrain the reachable outcomestokens that would break the schema are never emitted — but the schema constrains form, not intentmodel, possibly persuadedproposes next tokensCONSTRAINED DECODERtokens that would invalidatethe schema are never emittedalways a schema-valid objectno free text left to parseno extra fields, no trailing content, no format confusion for the consumer to misreadONE VALID OBJECT — FIELD BY FIELD{action: "refund"enum — bounded by the schemacustomer_id: "C-1042"pattern-validated referenceamount: 4900valid value — still attacker-influenceablenote: "Please expedite…"free text — an unconstrained channel}VALID IS NOT SAFEa transfer object can be schema-perfectwith the wrong recipient and the wrongamount — chosen because a documenttold the model to choose themthe policy check still has to happenPERMISSIVE SCHEMAS GIVE THE GROUND BACKidentifier typed as loose stringinstead of a validated patternenum still lists an unuseddestructive optionadditionalProperties leftopen on the objectMODEL OUTPUT IS UNTRUSTED INPUT AT EVERY CONSUMERbrowser renderingencode; no automaticremote fetchshell · query · templatetreat exactly asuser inputidentifiersresolve + authoriseat the consumeranother agentdata, unless deliberatelydecided otherwiseSTRUCTURE IS NOT AUTHORISATION — SCHEMA-VALID CAN STILL BE THE WRONG ACTIONthe recurring failure: validated once at the model boundary, trusted by five components downstream
Constrained decoding removes format confusion and bounds what a persuaded model can attempt — but every value and free-text field still needs validation and authorisation at each consumer.

Constraining the Shape of What the Model Can Emit

Constrained decoding restricts generation so the output must conform to a schema or grammar: at each step, tokens that would make the result invalid are excluded, so the model cannot produce malformed structure even if it would otherwise. As a security control this matters because it removes an entire class of downstream failure. When output is guaranteed to be a valid object with typed fields, the consumer does not need to parse free text, and an injected instruction cannot introduce structure the consumer will misread — no extra fields, no trailing content, no format confusion. It also shrinks the expressive room available to a persuaded model: if a tool takes an enumerated action and a validated identifier, the space of harmful calls is bounded by the schema regardless of what the model was persuaded to attempt. Constrain the shape, and you constrain the reachable outcomes.

  • Generation is restricted to tokens that keep the output schema-valid
  • Removes format confusion and parsing ambiguity at the consumer
  • Enumerated actions and validated identifiers bound what a persuaded model can attempt
  • Constraining shape constrains the set of reachable outcomes

The Limit: Valid Is Not Safe

The essential caveat is that a schema constrains form, not intent. A transfer object with a valid recipient identifier and a valid amount is schema-perfect and may still be the wrong recipient and the wrong amount, chosen because a document told the model to choose them. Every free-text field inside a valid object is an unconstrained channel — a message body, a note, a filename, a search query — and if any consumer renders or executes that field, the constraint has bought nothing there. Schemas are also frequently permissive in practice: an identifier typed as a string rather than a pattern-validated reference, an enum that includes an unused destructive option, an object with an open additional-properties policy. Constrained decoding is a strong hygiene control that eliminates a class of parsing failures. It is not authorisation, and treating it as one is a common overreach.

  • Schemas constrain form; the values inside remain attacker-influenceable
  • Free-text fields inside a valid object are unconstrained channels
  • Permissive schemas — loose string types, unused destructive enum values — give back the ground
  • Structure is not authorisation; the policy check still has to happen

Validate at Every Consumer

The complementary control is proper output handling wherever model output is used, and it follows ordinary application security practice applied to a source many teams forget to treat as untrusted. Text rendered in a browser needs encoding and a policy that prevents automatic fetching of remote resources referenced in it. Content passed to a shell, a query engine, a template renderer, or a deserialiser needs the same treatment as any user-supplied input, because that is exactly what it is. Identifiers should be resolved and authorised at the consumer rather than trusted because they arrived in a typed field. And output consumed by another agent should be treated as data unless you have deliberately decided otherwise. The recurring failure is a pipeline where output was validated once at the model boundary and then treated as internal by five components downstream.

  • Model output is untrusted input at every consumer, not only at the first one
  • Encode for the rendering context and disable automatic remote resource fetching
  • Resolve and authorise identifiers at the consumer, not on arrival
  • Validate once at the boundary and trust everywhere after is the pattern to avoid
  • OWASP's LLM Top 10 names this downstream failure improper output handling — treat every consumer as the enforcement point

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.