The AI Learning Hub Journal
◆ Mechanisms

Prompt Injection at Engineering Depth

One channel, two kinds of content — that is the whole bug classthe separation a parameterised query enforces structurally has no equivalent inside a promptPARAMETERISED QUERY — STRUCTURALSELECT … WHERE id = ?value: attacker textthe value can never become the statement —separation enforced by the interface, every timea boundary a diagram may rely onLLM PROMPT — ONE TOKEN SEQUENCEsystem promptuser msgretrieved doctool resultone undifferentiated sequence — the model inferseach part's role from learned patternsinstruction hierarchy: probabilistic preference, not a parser ruleWHY FILTERING CANNOT BE THE BOUNDARYunbounded, semantic expressionno finite pattern set covers every phrasingencoding routes around itbase64, translation, chunk reassembly, imagesthe attacker adaptsthey observe what passes; the filter lagskeep filters as cost-raising noise reduction — never as the thing the design depends onTHE RESPONSES THAT COUNT ARE ARCHITECTURALshrink capabilityfewer tools, no standing credsseparate rolesreader never holds privilegevalidate at consumersstructural, not inspectionreconstructable tracesscope what gets throughDESIGN AS IF INJECTION SUCCEEDS — THE MODEL DOES ONLY WHAT THE RUNTIME PERMITSonly structural controls belong on your diagram as boundaries; statistical ones are noise reduction
Injection is a structural property of one channel carrying instructions and data together — patch nothing, design as if it succeeds.

Why the Model Cannot Separate Instruction from Data

The separation that protects a parameterised database query is structural: the value travels in a different slot from the statement, and the parser never treats one as the other. A language model has no equivalent slot. Everything — system prompt, conversation, retrieved documents, tool results — becomes one token sequence, and the model infers the role of each part from patterns it learned during training. That inference is often good and never guaranteed. Providers train an instruction hierarchy so system content is weighted above user content, and it helps measurably, but it is a learned preference expressed through probabilities rather than an enforced rule, and text that convincingly resembles higher-priority content can shift the outcome. Understanding this precisely matters because it tells you which defences are structural and which are statistical, and only the structural ones belong on a diagram as a boundary.

  • No structural slot separates instruction from data; everything is one token sequence
  • Instruction hierarchy is learned and probabilistic, not parser-enforced
  • The model infers roles from patterns — convincing text can shift that inference
  • Distinguish structural controls from statistical ones when drawing boundaries
  • On the standard maps: this is OWASP LLM01 (prompt injection) — useful shorthand in design reviews and vendor conversations

Why Input Filtering Alone Fails

Filtering hostile instructions out of input is the first idea every team has, and it fails for reasons worth stating precisely rather than dismissing. The space of expressions that convey an instruction is unbounded and semantic: there is no finite pattern set covering every way to say what you want the model to do, across languages, indirection, politeness, and framing. Encoding widens the gap further, since content can arrive base64-encoded, translated, split across chunks that are reassembled in context, or embedded in an image or document the model reads. And the filter faces an adaptive attacker who can observe which phrasings pass. A semantic classifier does better than pattern matching, but it is a model with its own error rate operating on the same undecidable question. Filtering is a useful noise reducer that raises attacker cost. It is not a boundary, and no volume of tuning turns it into one.

  • Instruction expression is unbounded and semantic — no finite pattern set covers it
  • Encoding, translation, chunk reassembly and multimodal delivery route around surface filters
  • The attacker adapts by observing what passes; the filter cannot adapt as fast
  • Treat filtering as cost-raising noise reduction, never as a trust boundary

Direct Injection and What It Is Actually For

Direct injection — the user themselves supplying instructions intended to override the system prompt — is often dismissed as low severity on the grounds that users can only harm themselves. That reasoning holds only in a single-tenant, single-user context, and it breaks in several common cases. If the system prompt encodes business logic such as pricing rules, eligibility criteria, or discount authority, overriding it is a business-logic bypass. If the user's session can write to shared state — a corpus, a memory store, a ticket queue, a summary another person reads — the effect crosses to other users. If the assistant's output is trusted downstream as though it were system-generated, an overridden model becomes a way to inject content into a trusted channel. And system prompt extraction is a genuine confidentiality issue when the prompt contains internal policy, thresholds, or the tool inventory an attacker would otherwise have to guess.

  • Low severity only when the blast radius genuinely ends at the acting user
  • System prompts carrying business rules make override a business-logic bypass
  • Any write to shared state carries the effect across the user boundary
  • Prompt extraction leaks policy, thresholds, and the tool inventory

What Actually Reduces Risk

Because the mechanism is architectural, the effective responses are architectural too. Reduce what a persuaded model can do: fewer tools, narrower scopes, no standing credentials, no irreversible action without an independent check. Separate roles so the component exposed to untrusted content is not the component holding privilege — a planner that never sees retrieved text and an executor that cannot alter the plan is stronger than any filter, because the constraint is enforced outside the model. Validate output structurally where it is consumed, so a rendered link, a parsed command, or a downstream parser cannot be steered by free text. And maintain the ability to reconstruct any decision, so that when something does get through you can scope it. Filtering and delimiting still have a place on top of all this; they simply cannot be the thing the design depends on.

  • Shrink capability first — the persuaded model can only do what the runtime permits
  • Separate the component that reads untrusted content from the one holding privilege
  • Validate at every consumer of model output, structurally rather than by inspection
  • Keep reconstructable traces so a success can be scoped rather than guessed at

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.