Securing AI Systems
Threat-model, test, harden and respond — defending AI and agentic systems in practice
Read it in the library →Threat Modelling for AI Systems
Where classic threat models break on AI, how to inventory the real attack surface of a feature you own, drawing trust boundaries when the model itself consumes untrusted input, data-flow diagramming an LLM feature, enumerating techniques with MITRE ATLAS, and writing down the risks you have decided not to defend against.
The Attack Surface in Depth
Prompt injection at engineering depth and why filtering alone cannot hold, indirect injection through every ingestion channel, the lethal trifecta used as a design test, tool poisoning and the MCP supply chain, memory and context contamination, excessive agency, and the full inventory of exfiltration paths.
Hardening and Controls
Least privilege for agents and scoped tool access, sandboxing and egress control, approval gates on irreversible actions, structured output and constrained decoding, guardrail models and their false-positive cost, agent identity and full-trace audit logging, and defence in depth with an honest account of what no single control can hold.
Testing, Response, and Governance
Running an authorised adversarial exercise end to end, turning findings into security regression tests in CI, using automated adversarial testing without overstating it, detecting an AI incident in production, responding when the failing component is a model rather than a host, and closing the loop through post-incident review and governance.
Every lesson is free, with no sign-up. Reading happens in the library, where your progress is saved on your device.