The AI Learning Hub Journal

Agent-Specific Attack Surface

AgentPrompt Injectionvia retrieved dataTool Poisoningmalicious MCP serverConfused Deputymisuse of privilegesContext Contaminationpersistent memory exploitLateral Movementagent-to-agent escalation
The agentic attack surface — every connection is a potential vector

The New Threats

Confused-deputy attacks (agent uses its privileges on attacker behalf), tool poisoning (malicious MCP servers), context contamination (persistent memory exploit), lateral movement through agent tool chains. This is the frontier risk surface.

Tool Poisoning and the MCP Supply Chain

MCP servers are the new software supply chain. Known patterns: malicious tool descriptions that carry hidden instructions the model reads but the human never sees; "rug pulls" where a benign server updates itself into a malicious one after gaining adoption; typosquatted servers imitating popular ones; and over-scoped servers that request far more permissions than their function needs. Mitigations map to classic supply chain discipline: an allowlisted internal registry, version pinning, description auditing, and gateway-level policy on what each server may touch. This is precisely the problem agent registries and agent gateways exist to solve.

Memory and Context Contamination

Persistent agent memory turns a one-shot injection into a standing compromise: a poisoned instruction stored in memory re-executes across future sessions, long after the original malicious content is gone. Treat agent memory like a database with untrusted writes — validate what gets stored, scope memory per task where possible, expire aggressively, and make memory contents inspectable. Question worth asking of any agent deployment: "Can we inspect what our agents currently hold in memory — and could we tell if something malicious was written there last month?"

Other Attacks on AI Systems

Jailbreaks: bypass model safety alignment. Model poisoning: corrupt training data to install backdoors. Evasion: craft inputs that fool deployed models. Extraction: query a model to steal capabilities or training data. All concerns for organizations training or deploying their own models.

Defending the Agent Surface

The defense pattern is least privilege applied to a new actor class: scoped subagents that hold only the tools their task needs, approval gates (hooks) on irreversible actions, cryptographic agent identity so every action is attributable, and full-trace audit logging. None of this is exotic — it is IAM discipline extended to autonomous software. The organisations that get this right treat agents like employees: onboarding, scoped access, monitoring, and offboarding.

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.