Detecting an AI Incident in Production
What an AI Incident Looks Like
AI incidents often present without any of the signals conventional detection is built around. There is no malware, no anomalous process, no credential brute force, and frequently no failed request at all — every call is authenticated, authorised and successful, because the agent was permitted to do what it did. The visible symptom is usually behavioural: a user reports an answer containing data they should not have seen; a record was modified nobody can account for; a message was sent that no one wrote; an agent contacted a destination it has no reason to contact. Because the mechanism is persuasion rather than intrusion, detection has to sit on the semantics of what the system did, which means it is built on your own traces rather than on host or network telemetry. This is the practical reason trace quality is a security requirement rather than an engineering nicety.
- No malware, no failed auth — the actions are permitted and successful
- Symptoms are behavioural: wrong data returned, unexplained changes, unexpected messages
- Detection sits on application traces, not host or network telemetry
- Trace completeness is the limiting factor on whether detection is possible at all
Signals Worth Building
A small set of signals covers most of the ground. A privileged or irreversible tool call occurring in the same run as ingestion of externally sourced content is the highest-value single detection, because it directly encodes the dangerous sequence. New outbound destinations per agent, alerted on first appearance rather than on volume. Canary records and documents that no legitimate query should retrieve or transmit, which produce unambiguous alerts at almost no cost. Retrieval patterns unrelated to the user's request, which indicate the corpus is steering the session. Sharp movements in refusal rate, error rate, gate approval rate or run length, which frequently move before anyone notices a behavioural problem. Cross-user anomalies, such as one user's content appearing in another's session. Each needs an owner, a documented response, and a tuned threshold, or it becomes another ignored dashboard.
- Privileged action after external ingestion in one run — the highest-value signal
- Alert on first appearance of a new outbound destination per agent
- Canaries are cheap, unambiguous, and rarely deployed
- Rate shifts in refusals, errors, approvals and run length often move first
- MITRE ATLAS's exfiltration and impact tactics make a useful checklist when enumerating these behavioural signals
Where the Report Comes From
Plan for the likely case that the first indication arrives from a person rather than a monitor. A user notices an odd answer, a customer asks why they received a message, a partner reports content they should not have, or a researcher contacts you about a behaviour they found. That means two things need to exist in advance. A route by which behavioural reports about AI features reach the security team rather than terminating in a product support queue as a quality complaint — this is the most common way real incidents are lost. And a triage standard that distinguishes a model producing a poor answer from a model producing an answer that indicates influence or unauthorised access, because most reports are the former and the ones that are not look identical at first glance. Write examples of both into the triage guidance.
- Assume the first signal is a human report, not an alert
- Route behavioural reports about AI features to security, not only to support
- Distinguish poor-quality output from evidence of influence or unauthorised access
- Publish a channel for external researchers before you need one
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.