The AI Learning Hub Journal

Post-Incident Review and Governance

The review's output is four updated artefacts and a test"the model was tricked" names a property of the technology — control questions name a fixable causeTHE UNPRODUCTIVE REVIEW"the model was tricked""improve the prompt"unverifiable action — repeat finding laterCONTROL QUESTIONS THAT YIELD OWNED CHANGES· why did the agent have that capability at that moment?· why did that content reach the context — source reviewed?· which control was expected to catch it — absent, misconfigured, bypassed, or out of scope?· why did detection take that long — which signal was missing?· was it in the threat model — was it knowingly accepted?EVERY INCIDENT UPDATES FIVE ARTEFACTS — FEWER, AND THE REVIEW LEAKED ITS LESSONSthe incident, reviewedwith control questionsTHREAT MODEL — add the scenario, or correct the discounted likelihoodACCEPTED-RISK REGISTER — reopen any acceptance built on a disproven assumptionCONTROL SET — the change at the containing layer, with an owner and a dateREGRESSION SUITE — a test asserting the condition; recurrence is a build failureDETECTION — build the missing signal while the event's shape is still concreteGOVERNANCE THAT OPERATES — AND METRICS THAT DESCRIBE POSTURE, NOT ACTIVITYinventory per AI systemowner · purpose · data classes · tools · last reviewreview triggered by changea new tool, source, scope, or agent — proportionatenamed accountabilityfor systems, controls, and report routingmeasure: threat-model coverage · trifecta components trending down · time-to-detect from first trace evidence · regression coverageavoid: attempt counts · blocked-request totals · policies published — they move independently of whether anything is saferIF THE REVIEW PRODUCED NO TEST, NOTHING CLOSED — IF GOVERNANCE CANNOT ANSWER, IT IS PAPERfeed the threat model, the register, the controls, the suite and the detection layer — then measure posture
A useful review asks which control was expected to catch the incident and what it actually did, then updates the threat model, the risk register, the controls, the regression suite, and detection.

Reviewing Without Blaming the Model

The unproductive post-incident review for an AI failure concludes that the model was tricked and resolves to improve the prompt. It is unproductive because it identifies a property of the technology as the root cause and produces an action nobody can verify. A useful review asks control questions instead. Why did the agent have that capability at that moment? Why did that content reach the context, and was its source reviewed? Which control was expected to catch this, and what did it actually do — was it absent, misconfigured, bypassed, or working as designed against a case nobody considered? Why did detection take as long as it did, and which signal would have shortened it? Was this scenario in the threat model, and if so, was it in the accepted-risk register? Each of those questions yields a change with an owner; the prompt conclusion does not.

  • Model was tricked is a property, not a root cause
  • Ask which control was expected to catch it and what it actually did
  • Absent, misconfigured, bypassed, or out of scope — these lead to different fixes
  • Check whether the scenario was modelled and whether it was knowingly accepted

Feeding Findings Back

Every incident should update four artefacts, and a review that updates fewer has leaked its own lessons. The threat model gains the scenario if it was missing, or a corrected likelihood if it was present and discounted. The accepted-risk register is revisited, since an acceptance whose assumption has now been demonstrated false must be reopened rather than quietly retained. The control set gains the specific change, at the layer that contains the issue, with an owner and a date. And the regression suite gains a test asserting the condition, so the same path failing again is a build failure rather than a second incident. Also feed the detection layer: if the signal that would have caught this did not exist, build it now, while the shape of the event is still clear enough to specify a threshold worth setting.

  • Update the threat model, the accepted-risk register, the controls, and the test suite
  • An acceptance built on a disproven assumption must be reopened, not retained
  • Add the missing detection signal while the event shape is still concrete
  • A review that produces no test has not closed anything

Governance That Is Actually Operable

A workable governance layer needs three things and not much more. An inventory of AI systems in use, with owner, purpose, data classes, tool and permission list, and the date of the last threat model — most organisations cannot produce this, and its absence is why nobody can answer basic questions during an incident. A gate at the point of change, so adding a tool, a data source, a permission scope, or a new agent triggers a proportionate review rather than a full re-assessment. And clear ownership: someone accountable for each system, someone accountable for the controls, and a defined route for behavioural reports. Everything beyond this — policy documents, committees, generic questionnaires — tends to consume effort without changing what the systems can do. Governance that cannot answer what our agents can reach today is documentation, not control.

  • Inventory with owner, purpose, data classes, tools, permissions, last review date
  • Proportionate review triggered by change, not periodic full re-assessment
  • Named accountability for systems, controls, and behavioural report routing
  • If it cannot answer what agents can reach today, it is not governance

Metrics That Reflect Reality

Choose metrics that describe posture rather than activity, because activity metrics reward motion. Coverage: the proportion of deployed AI features with a current threat model and an enumerated tool and permission inventory. Containment: how many components still combine private data, untrusted content and an outbound channel, tracked as a number that should fall. Time to detect and time to contain for AI-specific incidents, measured from the earliest evidence in traces rather than from the report. Regression coverage: proportion of accepted findings with a test that would fail if the condition returned. Gate quality: approval rates and decision times, which reveal whether gates are being read. Avoid attempt counts, blocked-request totals, and policy completion rates — they move independently of whether anything is safer, and they are the metrics most likely to be requested.

  • Coverage of threat models and permission inventories across deployed features
  • Count of components still holding all three trifecta properties, trending down
  • Time to detect measured from earliest trace evidence, not from the report
  • Avoid attempt counts and blocked-request totals — motion, not posture

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.