The AI Learning Hub Journal

Trust Boundaries When the Model Is Untrusted

Draw the boundary around the model — not behind itthe model processes untrusted input, so it produces untrusted output — every consumer sits across a boundaryTHE INTUITIVE PLACEMENT — WRONG FOR AN LLM FEATUREuser inputvalidatethe only boundary drawnassumed trustedmodelrenderer · parser · shell · DB — uncheckedthis placement implies model output can be trusted by whatever consumes itDRAW IT AROUND THE MODEL — OUTPUT IS UNTRUSTED, EVERY TIMEuser input (authenticated)retrieved content (mixed)tool results (untrusted)UNTRUSTED ZONEcontext window+ modeluntrusted outputrenderersanitise like internet inputparser · shell · DBvalidate before anything executestool dispatcheran authorisation decision, not a callevery arrow leaving the model crosses a trust boundary — treat each as input arriving from the internetPROVENANCE MAKES THE BOUNDARY ENFORCEABLE — BY THE RUNTIME, NOT THE MODELtag content by source, keep tags through summarisation and handoffs, and gate privileged tools on what entered the turnMODEL OUTPUT IS UNTRUSTED — NO MATTER HOW TRUSTED THE INPUT LOOKEDevery agent-to-agent edge is a boundary too, and delegation must never silently upgrade authority
Place the trust boundary around the model: its output is untrusted by every consumer, and a tool call is an authorisation decision.

Draw the Boundary Around the Model, Not Behind It

The single most consequential drawing decision is where the trust boundary sits relative to the model. The intuitive placement puts the boundary at the API edge, treats user input as untrusted, and treats everything after validation as internal. That placement is wrong for an LLM feature, because it implies the model's output can be trusted by whatever consumes it. The correct placement treats the model as a component that processes untrusted input and therefore produces untrusted output — every time. Model output crossing into a renderer, a parser, a shell, a database, or another agent crosses a trust boundary and needs the same treatment you would give input arriving from the internet. This one change reframes most downstream design: output handling stops being formatting and becomes validation, and a tool call stops being an internal function invocation and becomes an authorisation decision.

  • Model output is untrusted output, regardless of how trusted the input looked
  • Every consumer of model output sits across a trust boundary from it
  • Tool invocation is an authorisation decision, not an internal function call
  • Improper output handling is a distinct failure class from injection — treat it separately
  • On the standard maps: OWASP's LLM Top 10 carries improper output handling as its own entry, separate from prompt injection

Provenance as a First-Class Property

Once you accept that the context window mixes content of different trust levels, the practical question becomes whether the system can tell them apart at the point where it matters. Provenance means tagging content with where it came from — system-authored, authenticated user, retrieved from an internal reviewed corpus, retrieved from the open web, returned by a third-party tool — and carrying that tag through every transformation. The value is not that the model reliably respects the tags; it does not. The value is that the surrounding runtime can. A policy that says a privileged tool may not be called during a turn in which open-web content entered the context is enforceable in code, and it is only enforceable if provenance survived summarisation, chunking, caching, and the hop between agent and subagent. Test that survival explicitly, because losing tags in a transformation is the common failure.

  • Tag content by source and carry the tag through every hop and transformation
  • The runtime enforces provenance policy — the model is not the enforcement point
  • Summarisation, chunking, and subagent handoffs are where tags are commonly dropped
  • Provenance enables rules like: no privileged action in a turn that ingested untrusted content

Boundaries Between Agents

Multi-agent designs create trust relationships that are almost never written down. When an orchestrator delegates to a subagent and receives text back, that text is input from a component that was itself exposed to untrusted content — so the boundary between them is real even though both are yours. The same holds for a peer agent operated by another team, and much more strongly for one operated by another organisation, where you can see the interface but not the reasoning, the tools, or the data behind it. Draw a boundary at every agent-to-agent edge and decide explicitly what crosses it: results only, or results plus instructions the receiver will act on. Then decide what authority the receiver applies — its own, or the originating user's. Delegation that silently upgrades authority is the multi-agent form of the confused deputy, and it is easy to build by accident.

  • Every agent-to-agent edge is a trust boundary, including between your own components
  • A peer agent is opaque by design: you see results, never its context or tools
  • Decide whether delegated output is treated as data or as instruction — write it down
  • Delegation must not upgrade authority; carry the originating user's scope through the chain

Humans Inside the Boundary

When a design includes a reviewer or approver, that person is part of the control flow and belongs on the diagram as a component with inputs and failure modes. Their input is a summary produced by the system being attacked, which means the attacker can influence what the approver reads. Their failure modes are well documented: fatigue when volume is high, automation bias when the system is usually right, and misreading when the summary is plausible but incomplete. Model the human path as you would any other: what information reaches the decision point, whether it is attacker-influenceable, how long the decision takes, and what the default is when the reviewer does nothing. A gate whose default is to proceed after a timeout is not a gate. A gate whose summary is written by the compromised component is a gate the attacker helps design.

  • Approvers are components: give them inputs, failure modes, and a place on the diagram
  • The approval summary is model output — attacker-influenceable by construction
  • Automation bias and fatigue are predictable, so design for volume, not for vigilance
  • A timeout that defaults to proceed converts the gate into a delay

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.