The AI Learning Hub Journal

Security Attack Surface of Agents

AgentPrompt Injectionvia retrieved dataTool Poisoningmalicious MCP serverConfused Deputymisuse of privilegesContext Contaminationpersistent memory exploitLateral Movementagent-to-agent escalation
The agentic attack surface — every connection is a potential vector

Prompt Injection Grows Teeth

In a chatbot, prompt injection produces a bad answer. In an agent, it produces bad actions — the same loop that lets the model act on the world lets an attacker act through it. The critical insight is that the injection rarely arrives from the user: it arrives through the observe step. Any content an agent reads — a webpage, an email, a ticket, a repository README, a tool result — is a potential instruction channel, because models cannot reliably distinguish "data I was given to process" from "instructions I should follow." The canonical danger pattern is the combination of three properties: access to private data, exposure to untrusted content, and the ability to communicate externally. An agent holding all three can be turned into an exfiltration engine by a single well-placed paragraph. Remove or gate at least one leg, because no prompt-level defence reliably holds on its own.

  • Direct injection comes from the user; indirect injection rides in on anything the agent reads
  • Every tool result is untrusted input re-entering the reasoning context
  • Danger trifecta: private data + untrusted content + external communication in one agent
  • "Ignore previous instructions" is the toy version — real payloads are contextual and polite

Confused Deputy and Tool Poisoning

The confused deputy is a classic security failure with a sharp agentic edge: a privileged intermediary tricked into using its authority on an attacker's behalf. An agent is a deputy by construction — it runs with the user's credentials and permissions, and any injected instruction executes with that borrowed authority. The attacker never touches your API keys; they persuade the entity that legitimately holds them. Tool poisoning attacks the supply chain instead: because tool descriptions are consumed by the model as trusted context, a malicious or compromised MCP server can embed hidden directives in its own metadata — instructions invisible in a casual review but followed by the model. Variants include rug pulls, where a server behaves impeccably until wide adoption and then swaps its definitions mid-session, and cross-tool shadowing, where one hostile tool's description manipulates how the agent uses other, honest tools.

  • The agent's legitimate authority is the weapon — attacks borrow it rather than steal it
  • Tool descriptions are prompt content from a third party: audit them like executable code
  • Rug pull: trusted server changes behaviour after adoption — pin and review versions
  • One poisoned tool can corrupt an agent's use of every other tool in the session

Memory Contamination

Persistence turns a transient attack into an infection. Agents increasingly carry state across sessions — scratchpad files, memory stores, retrieval indexes, learned preferences — and each store is writable, directly or indirectly, by content the agent processed. An injected instruction that says remember this becomes a standing order: the malicious payload is gone from the context, but its residue re-enters every future session as trusted memory, long after the hostile document is deleted. Retrieval-backed memory widens the door — poisoning the corpus an agent retrieves from plants payloads that surface on the attacker's chosen topic. This breaks the comfortable assumption that each session starts clean, and it makes provenance a first-class requirement: memory written under the influence of untrusted content is itself untrusted, and an agent's memory should be inspectable, attributable, and revocable like any other privileged store.

  • Session-scoped injection ends with the session; memory-scoped injection compounds
  • Track provenance: what wrote each memory, and what content was in context when it did
  • Retrieval corpora are memory too — poisoned documents are time-delayed injections
  • Provide a working "forget" path: contaminated memory you cannot excise is a standing backdoor

Defence: Least Privilege, Sandboxes, Gates

No current technique makes a model reliably immune to injection, so effective defence assumes compromise and constrains blast radius — a familiar posture: zero trust applied to a new kind of insider. Least privilege comes first: an agent gets the narrowest toolset and scopes the task allows, read-only where read-only suffices, per-task credentials rather than standing ones. Sandboxing contains execution — code and file operations inside containers with explicit filesystem scope and default-deny egress, so even a fully hijacked agent cannot reach what the sandbox excludes. Approval gates put a human decision on the irreversible subset: sending, deleting, paying, deploying. Gate by consequence, not by frequency — a gate on everything trains reflexive clicking, which is no gate at all. Layer on top: strip or neutralise suspicious tool-result content, monitor for anomalous action sequences, and rehearse the incident path for the day a gate is talked through.

  • Assume the model can be turned; engineer so a turned model has nowhere to go
  • Per-task scoped credentials beat standing broad ones — expiry is a security control
  • Sandbox with default-deny egress: exfiltration needs a channel, so refuse to provide one
  • Reserve human approval for irreversible actions; approval fatigue is a vulnerability class
  • Securing AI Systems goes further: threat modelling, the jailbreak taxonomy, hardening controls, and running an authorised red-team exercise

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.