Security Attack Surface of Agents
Prompt Injection Grows Teeth
In a chatbot, prompt injection produces a bad answer. In an agent, it produces bad actions — the same loop that lets the model act on the world lets an attacker act through it. The critical insight is that the injection rarely arrives from the user: it arrives through the observe step. Any content an agent reads — a webpage, an email, a ticket, a repository README, a tool result — is a potential instruction channel, because models cannot reliably distinguish "data I was given to process" from "instructions I should follow." The canonical danger pattern is the combination of three properties: access to private data, exposure to untrusted content, and the ability to communicate externally. An agent holding all three can be turned into an exfiltration engine by a single well-placed paragraph. Remove or gate at least one leg, because no prompt-level defence reliably holds on its own.
- Direct injection comes from the user; indirect injection rides in on anything the agent reads
- Every tool result is untrusted input re-entering the reasoning context
- Danger trifecta: private data + untrusted content + external communication in one agent
- "Ignore previous instructions" is the toy version — real payloads are contextual and polite
Confused Deputy and Tool Poisoning
The confused deputy is a classic security failure with a sharp agentic edge: a privileged intermediary tricked into using its authority on an attacker's behalf. An agent is a deputy by construction — it runs with the user's credentials and permissions, and any injected instruction executes with that borrowed authority. The attacker never touches your API keys; they persuade the entity that legitimately holds them. Tool poisoning attacks the supply chain instead: because tool descriptions are consumed by the model as trusted context, a malicious or compromised MCP server can embed hidden directives in its own metadata — instructions invisible in a casual review but followed by the model. Variants include rug pulls, where a server behaves impeccably until wide adoption and then swaps its definitions mid-session, and cross-tool shadowing, where one hostile tool's description manipulates how the agent uses other, honest tools.
- The agent's legitimate authority is the weapon — attacks borrow it rather than steal it
- Tool descriptions are prompt content from a third party: audit them like executable code
- Rug pull: trusted server changes behaviour after adoption — pin and review versions
- One poisoned tool can corrupt an agent's use of every other tool in the session
Memory Contamination
Persistence turns a transient attack into an infection. Agents increasingly carry state across sessions — scratchpad files, memory stores, retrieval indexes, learned preferences — and each store is writable, directly or indirectly, by content the agent processed. An injected instruction that says remember this becomes a standing order: the malicious payload is gone from the context, but its residue re-enters every future session as trusted memory, long after the hostile document is deleted. Retrieval-backed memory widens the door — poisoning the corpus an agent retrieves from plants payloads that surface on the attacker's chosen topic. This breaks the comfortable assumption that each session starts clean, and it makes provenance a first-class requirement: memory written under the influence of untrusted content is itself untrusted, and an agent's memory should be inspectable, attributable, and revocable like any other privileged store.
- Session-scoped injection ends with the session; memory-scoped injection compounds
- Track provenance: what wrote each memory, and what content was in context when it did
- Retrieval corpora are memory too — poisoned documents are time-delayed injections
- Provide a working "forget" path: contaminated memory you cannot excise is a standing backdoor
Defence: Least Privilege, Sandboxes, Gates
No current technique makes a model reliably immune to injection, so effective defence assumes compromise and constrains blast radius — a familiar posture: zero trust applied to a new kind of insider. Least privilege comes first: an agent gets the narrowest toolset and scopes the task allows, read-only where read-only suffices, per-task credentials rather than standing ones. Sandboxing contains execution — code and file operations inside containers with explicit filesystem scope and default-deny egress, so even a fully hijacked agent cannot reach what the sandbox excludes. Approval gates put a human decision on the irreversible subset: sending, deleting, paying, deploying. Gate by consequence, not by frequency — a gate on everything trains reflexive clicking, which is no gate at all. Layer on top: strip or neutralise suspicious tool-result content, monitor for anomalous action sequences, and rehearse the incident path for the day a gate is talked through.
- Assume the model can be turned; engineer so a turned model has nowhere to go
- Per-task scoped credentials beat standing broad ones — expiry is a security control
- Sandbox with default-deny egress: exfiltration needs a channel, so refuse to provide one
- Reserve human approval for irreversible actions; approval fatigue is a vulnerability class
- Securing AI Systems goes further: threat modelling, the jailbreak taxonomy, hardening controls, and running an authorised red-team exercise
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.