The AI Learning Hub Journal

Detecting an AI Incident in Production

Nothing failed, and that is the problemthe mechanism is persuasion, not intrusion — every call is authenticated, authorised, and successfulCLASSIC IOCs — ABSENT✕ no malware, no anomalous process✕ no credential brute force✕ no failed or denied requests✕ nothing host or network telemetry flagsthe agent was permitted to do what it didWHAT ACTUALLY SURFACES — BEHAVIOUR· an answer contains data the user should not see· a record modified nobody can account for· a message sent that no one wrote· a destination contacted for no reasonsemantic symptoms, visible only in your tracesSIGNALS WORTH BUILDING — SEQUENCE AND DESTINATION, NOT CONTENTprivileged call after externalingestion, same run — highest valuefirst-seen outbounddestination, per agentcanary records that shouldnever move — cheap, unambiguousretrieval unrelated to theuser’s requestrate shifts: refusals, errors,approvals, run lengthcross-user contentappearing in a sessioneach signal needs an owner, a documented response, and a tuned threshold — or it is another unread dashboardPLAN FOR THE FIRST SIGNAL BEING A PERSONa user, a customer, a partneror a researcher reportsodd behaviourproduct support queue — logged asa quality complaint; incident lostsecurity triage — trained to tell apoor answer from signs of influencethe most common way realincidents are lostpublish the researcher channelbefore you need oneDETECTION SITS ON YOUR TRACES — TRACE COMPLETENESS DECIDES IF IT IS POSSIBLEroute behavioural reports about AI features to security, not into a support queue
AI incidents present as permitted, successful actions with behavioural symptoms, so detection lives in your traces — and the first signal is usually a person, who must reach security rather than a support queue.

What an AI Incident Looks Like

AI incidents often present without any of the signals conventional detection is built around. There is no malware, no anomalous process, no credential brute force, and frequently no failed request at all — every call is authenticated, authorised and successful, because the agent was permitted to do what it did. The visible symptom is usually behavioural: a user reports an answer containing data they should not have seen; a record was modified nobody can account for; a message was sent that no one wrote; an agent contacted a destination it has no reason to contact. Because the mechanism is persuasion rather than intrusion, detection has to sit on the semantics of what the system did, which means it is built on your own traces rather than on host or network telemetry. This is the practical reason trace quality is a security requirement rather than an engineering nicety.

  • No malware, no failed auth — the actions are permitted and successful
  • Symptoms are behavioural: wrong data returned, unexplained changes, unexpected messages
  • Detection sits on application traces, not host or network telemetry
  • Trace completeness is the limiting factor on whether detection is possible at all

Signals Worth Building

A small set of signals covers most of the ground. A privileged or irreversible tool call occurring in the same run as ingestion of externally sourced content is the highest-value single detection, because it directly encodes the dangerous sequence. New outbound destinations per agent, alerted on first appearance rather than on volume. Canary records and documents that no legitimate query should retrieve or transmit, which produce unambiguous alerts at almost no cost. Retrieval patterns unrelated to the user's request, which indicate the corpus is steering the session. Sharp movements in refusal rate, error rate, gate approval rate or run length, which frequently move before anyone notices a behavioural problem. Cross-user anomalies, such as one user's content appearing in another's session. Each needs an owner, a documented response, and a tuned threshold, or it becomes another ignored dashboard.

  • Privileged action after external ingestion in one run — the highest-value signal
  • Alert on first appearance of a new outbound destination per agent
  • Canaries are cheap, unambiguous, and rarely deployed
  • Rate shifts in refusals, errors, approvals and run length often move first
  • MITRE ATLAS's exfiltration and impact tactics make a useful checklist when enumerating these behavioural signals

Where the Report Comes From

Plan for the likely case that the first indication arrives from a person rather than a monitor. A user notices an odd answer, a customer asks why they received a message, a partner reports content they should not have, or a researcher contacts you about a behaviour they found. That means two things need to exist in advance. A route by which behavioural reports about AI features reach the security team rather than terminating in a product support queue as a quality complaint — this is the most common way real incidents are lost. And a triage standard that distinguishes a model producing a poor answer from a model producing an answer that indicates influence or unauthorised access, because most reports are the former and the ones that are not look identical at first glance. Write examples of both into the triage guidance.

  • Assume the first signal is a human report, not an alert
  • Route behavioural reports about AI features to security, not only to support
  • Distinguish poor-quality output from evidence of influence or unauthorised access
  • Publish a channel for external researchers before you need one

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.