Indirect Injection and the Ingestion Inventory
The Channel Inventory Is the Work
Indirect injection is the serious case because the attacker never touches your product: they place content where your system will read it, and the victim is whoever the system is serving at the time. Defending it starts with an inventory that is usually longer than teams expect. Retrieved documents and their metadata. Web pages fetched by a browsing tool. Email bodies, attachments, and headers. Ticket titles, descriptions and comments. Repository files, commit messages, code comments, dependency manifests, and CI output. Calendar invites and their notes. Filenames, alt text, and document properties. Tool responses from third-party APIs, including error strings. Content submitted by one user that reaches another user's session. Transcripts and captions from audio or video. Every one of these is an instruction channel because the model reads it, and any channel you have not listed is one you cannot have tested.
- The attacker never authenticates — they position content and wait
- Metadata counts: filenames, headers, alt text, document properties, error strings
- Third-party tool responses are attacker-influenceable if their upstream data is
- An unlisted channel is an untested channel — the inventory is the deliverable
- OWASP LLM01 (prompt injection) covers this indirect form explicitly; MITRE ATLAS catalogues it among its prompt injection techniques
Delayed and Conditional Activation
Two properties make indirect injection harder to test than a single-shot attack, and both should shape your test design. Delay: content is planted long before it is read, so the poisoned document sits in a corpus for weeks and the attack happens when some user asks a related question. There is no temporal correlation between the malicious action and any attacker activity, which defeats detection strategies built on session-level anomaly. Conditionality: planted content can be written to act only under specific circumstances — when a particular topic is retrieved, when a particular tool appears in the context, when the reader appears to hold elevated access. This defeats sampling-based review, because the content behaves benignly whenever it is examined and only activates in the situation it was written for. Assume both properties when designing corpus review and detection.
- Planting and activation are separated in time, breaking session-level correlation
- Conditional content stays benign under review and activates only in its target situation
- Sampling a corpus for hostile content is weak against conditional payloads
- Detection must sit at the effect — anomalous actions and egress — not only at ingestion
Second-Order and Cross-User Paths
Content frequently passes through several components before it reaches the model that acts, and each hop is a chance for trust to be laundered. A summariser reads a hostile page and writes a summary into a store; the summary is later read by an agent with tools, and it now looks like internal system-generated content. A support agent reads a customer message and writes a case note; the note is later read by a different agent with billing access. A code review agent reads a pull request and its comment becomes part of a build log another automation ingests. In each case the original provenance is lost at the first hop, which is exactly why provenance must survive transformation. When mapping these paths, follow the content rather than the request: ask where the text ends up, who reads it next, and with what privileges — not just who asked the original question.
- Each hop launders provenance unless the tag is carried deliberately
- Summaries and case notes are the classic laundering points
- Follow the content forward, not the request backward
- Ask what privileges the eventual reader holds, not the original requester
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.