The AI Learning Hub Journal

Prompt Injection: The Dominant Threat

Direct InjectionAttackerLLMLower enterprise risk(direct user manipulation)Indirect InjectionUser QueryLLMRetrieves docPoisoned contentHidden instructions in PDF, web page, emailLLM acts on attacker instructions
Indirect injection is the dominant practical threat — payload sits in retrieved data, not user input

Direct vs Indirect

Direct: attacker types malicious instructions into the prompt directly. Indirect: payload sits in retrieved data (documents, web pages, emails, tool outputs) that the model later reads. Indirect injection is the dominant practical threat in enterprise.

Why Indirect Is Worse

The user did nothing wrong. They just asked the agent to summarize a document — the document was poisoned. Defending requires treating all retrieved content as untrusted, including content from your own systems.

The Lethal Trifecta

A useful mental model for when injection becomes exfiltration: an agent is dangerous when it combines (1) access to private data, (2) exposure to untrusted content, and (3) an outbound channel — the ability to send data somewhere an attacker can read it. Remove any one leg and the worst outcome is contained. This is why "the agent reads email AND can browse the web AND has access to the CRM" should trigger an architecture review, not a feature celebration. Anyone who can whiteboard this trifecta will evaluate an agent deployment faster than anyone reading from a feature list.

Enterprise Mitigation Architecture

Layered defense: (1) Input sanitization — strip or escape instruction-like patterns before they reach the model context. (2) Privilege separation — the agent that reads external documents should not also have access to high-privilege tools like account management or data export. (3) Output validation — constrain agent outputs to structured formats; free-form tool calls from agent output are higher risk than schema-validated calls. (4) Model-as-judge — a second model or rule layer validates that agent actions match the original user intent. (5) Audit logging — full trace of retrieved content, model input, and tool calls so injections can be forensically reconstructed.

Questions to Ask Before Deployment

How do you authorize which MCP servers or tools your agents can call? When your agent reads an email or document, is that content treated as trusted or untrusted input? Do you log the full context window — not just the user query, but retrieved content and tool results — for agent sessions? These questions surface injection exposure before it becomes an incident.

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.