Prompt Injection at Engineering Depth
Why the Model Cannot Separate Instruction from Data
The separation that protects a parameterised database query is structural: the value travels in a different slot from the statement, and the parser never treats one as the other. A language model has no equivalent slot. Everything — system prompt, conversation, retrieved documents, tool results — becomes one token sequence, and the model infers the role of each part from patterns it learned during training. That inference is often good and never guaranteed. Providers train an instruction hierarchy so system content is weighted above user content, and it helps measurably, but it is a learned preference expressed through probabilities rather than an enforced rule, and text that convincingly resembles higher-priority content can shift the outcome. Understanding this precisely matters because it tells you which defences are structural and which are statistical, and only the structural ones belong on a diagram as a boundary.
- No structural slot separates instruction from data; everything is one token sequence
- Instruction hierarchy is learned and probabilistic, not parser-enforced
- The model infers roles from patterns — convincing text can shift that inference
- Distinguish structural controls from statistical ones when drawing boundaries
- On the standard maps: this is OWASP LLM01 (prompt injection) — useful shorthand in design reviews and vendor conversations
Why Input Filtering Alone Fails
Filtering hostile instructions out of input is the first idea every team has, and it fails for reasons worth stating precisely rather than dismissing. The space of expressions that convey an instruction is unbounded and semantic: there is no finite pattern set covering every way to say what you want the model to do, across languages, indirection, politeness, and framing. Encoding widens the gap further, since content can arrive base64-encoded, translated, split across chunks that are reassembled in context, or embedded in an image or document the model reads. And the filter faces an adaptive attacker who can observe which phrasings pass. A semantic classifier does better than pattern matching, but it is a model with its own error rate operating on the same undecidable question. Filtering is a useful noise reducer that raises attacker cost. It is not a boundary, and no volume of tuning turns it into one.
- Instruction expression is unbounded and semantic — no finite pattern set covers it
- Encoding, translation, chunk reassembly and multimodal delivery route around surface filters
- The attacker adapts by observing what passes; the filter cannot adapt as fast
- Treat filtering as cost-raising noise reduction, never as a trust boundary
Direct Injection and What It Is Actually For
Direct injection — the user themselves supplying instructions intended to override the system prompt — is often dismissed as low severity on the grounds that users can only harm themselves. That reasoning holds only in a single-tenant, single-user context, and it breaks in several common cases. If the system prompt encodes business logic such as pricing rules, eligibility criteria, or discount authority, overriding it is a business-logic bypass. If the user's session can write to shared state — a corpus, a memory store, a ticket queue, a summary another person reads — the effect crosses to other users. If the assistant's output is trusted downstream as though it were system-generated, an overridden model becomes a way to inject content into a trusted channel. And system prompt extraction is a genuine confidentiality issue when the prompt contains internal policy, thresholds, or the tool inventory an attacker would otherwise have to guess.
- Low severity only when the blast radius genuinely ends at the acting user
- System prompts carrying business rules make override a business-logic bypass
- Any write to shared state carries the effect across the user boundary
- Prompt extraction leaks policy, thresholds, and the tool inventory
What Actually Reduces Risk
Because the mechanism is architectural, the effective responses are architectural too. Reduce what a persuaded model can do: fewer tools, narrower scopes, no standing credentials, no irreversible action without an independent check. Separate roles so the component exposed to untrusted content is not the component holding privilege — a planner that never sees retrieved text and an executor that cannot alter the plan is stronger than any filter, because the constraint is enforced outside the model. Validate output structurally where it is consumed, so a rendered link, a parsed command, or a downstream parser cannot be steered by free text. And maintain the ability to reconstruct any decision, so that when something does get through you can scope it. Filtering and delimiting still have a place on top of all this; they simply cannot be the thing the design depends on.
- Shrink capability first — the persuaded model can only do what the runtime permits
- Separate the component that reads untrusted content from the one holding privilege
- Validate at every consumer of model output, structurally rather than by inspection
- Keep reconstructable traces so a success can be scoped rather than guessed at
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.