The AI Learning Hub Journal

Indirect Injection and the Ingestion Inventory

The attacker never logs in — they position content and waitdefence starts with an inventory of every path by which content reaches the modelINGESTION INVENTORY — LONGER THAN YOU EXPECTretrieved docs + their metadataweb pages a browsing tool fetchesemail bodies, attachments, headersticket titles, comments, fieldsrepo files, commits, CI outputcalendar invites and notesfilenames, alt text, doc properties3rd-party tool responses + errorsone user's content, another's sessionaudio / video transcripts, captionscontext window + modelacting for whoever thesystem serves right nowevery channel the model readsis an instruction channelPLANTED LONG BEFORE IT ACTS — AND ONLY WHEN CONDITIONS MATCHcontent plantedno product touchedweeks passreviews see benign textrelated question askedconditional payload activatesaction firesnothing to correlate withSECOND-ORDER PATHS LAUNDER TRUST AT EVERY HOPhostile pageexternal, untrustedsummariser reads itwrites a case notestored summarynow looks internalagent with billing toolsacts on the notefollow the content forward — who reads it next, and with what privileges?AN UNLISTED CHANNEL IS AN UNTESTED CHANNEL — THE INVENTORY IS THE DELIVERABLEsampling review misses conditional payloads — put detection at the effect: anomalous actions and egress
Inventory every path content takes into the model, assume delay and conditionality, and put detection at the effect rather than at ingestion.

The Channel Inventory Is the Work

Indirect injection is the serious case because the attacker never touches your product: they place content where your system will read it, and the victim is whoever the system is serving at the time. Defending it starts with an inventory that is usually longer than teams expect. Retrieved documents and their metadata. Web pages fetched by a browsing tool. Email bodies, attachments, and headers. Ticket titles, descriptions and comments. Repository files, commit messages, code comments, dependency manifests, and CI output. Calendar invites and their notes. Filenames, alt text, and document properties. Tool responses from third-party APIs, including error strings. Content submitted by one user that reaches another user's session. Transcripts and captions from audio or video. Every one of these is an instruction channel because the model reads it, and any channel you have not listed is one you cannot have tested.

  • The attacker never authenticates — they position content and wait
  • Metadata counts: filenames, headers, alt text, document properties, error strings
  • Third-party tool responses are attacker-influenceable if their upstream data is
  • An unlisted channel is an untested channel — the inventory is the deliverable
  • OWASP LLM01 (prompt injection) covers this indirect form explicitly; MITRE ATLAS catalogues it among its prompt injection techniques

Delayed and Conditional Activation

Two properties make indirect injection harder to test than a single-shot attack, and both should shape your test design. Delay: content is planted long before it is read, so the poisoned document sits in a corpus for weeks and the attack happens when some user asks a related question. There is no temporal correlation between the malicious action and any attacker activity, which defeats detection strategies built on session-level anomaly. Conditionality: planted content can be written to act only under specific circumstances — when a particular topic is retrieved, when a particular tool appears in the context, when the reader appears to hold elevated access. This defeats sampling-based review, because the content behaves benignly whenever it is examined and only activates in the situation it was written for. Assume both properties when designing corpus review and detection.

  • Planting and activation are separated in time, breaking session-level correlation
  • Conditional content stays benign under review and activates only in its target situation
  • Sampling a corpus for hostile content is weak against conditional payloads
  • Detection must sit at the effect — anomalous actions and egress — not only at ingestion

Second-Order and Cross-User Paths

Content frequently passes through several components before it reaches the model that acts, and each hop is a chance for trust to be laundered. A summariser reads a hostile page and writes a summary into a store; the summary is later read by an agent with tools, and it now looks like internal system-generated content. A support agent reads a customer message and writes a case note; the note is later read by a different agent with billing access. A code review agent reads a pull request and its comment becomes part of a build log another automation ingests. In each case the original provenance is lost at the first hop, which is exactly why provenance must survive transformation. When mapping these paths, follow the content rather than the request: ask where the text ends up, who reads it next, and with what privileges — not just who asked the original question.

  • Each hop launders provenance unless the tag is carried deliberately
  • Summaries and case notes are the classic laundering points
  • Follow the content forward, not the request backward
  • Ask what privileges the eventual reader holds, not the original requester

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.