The AI Learning Hub Journal

Deciding What You Will Not Defend

What you will not defend, written down and signedafter an incident, unstated acceptance is indistinguishable from oversight — the register is the differenceANATOMY OF ONE ACCEPTED-RISK ENTRYSCENARIOthe outcome the attacker achieves —not the technique they useREASONcost · low likelihood · mitigation wouldbreak the feature · covered elsewhereACCEPTED BYa named person, with a date —unsigned is unacceptedREOPEN WHENnew data class · new write tool · newuser population · industry incidentSTILL IN PLACEpartial measures — detected but notprevented is a different postureREVIEW ON TRIGGERS, NOT CALENDARa new tool landsa new data class entersthe user population changesthe model deployment changesre-derive the acceptanceaccepted because tools were read-only?void the moment a write tool landsWHERE PREVENTION IS UNAVAILABLE — SAY SO, THEN CONTAIN AND DETECTPROMPT INJECTIONno technique reliably prevents itCONTAINnarrower tools · denied egress · gatesDETECTodd tool sequences · canaries · rate shiftsINJECTION EXPECTED, CONTAINED, DETECTED — BEATS CLAIMING A FILTER PREVENTS ITa register that has never had an entry reopened is not being read
Write the scenario, the reason, the named accepter with a date, and the triggers that reopen the decision — unsigned is unaccepted.

Every Threat Model Has an Out-of-Scope List

A threat model without an explicit out-of-scope list is not complete, it is evasive. Every real system accepts risks: you may decide not to defend against a fully compromised model provider, a malicious insider with corpus write access and review authority, a determined attacker with the user's own credentials, or a state-level adversary targeting your specific feature. Those may all be correct decisions given cost, likelihood, and what the feature is worth. What is not acceptable is leaving them unstated, because then nobody can tell whether the risk was assessed and accepted or simply never considered — and after an incident that distinction is the entire conversation. Writing the list also has a useful forcing effect during design: teams that must name what they are not defending against usually discover one item on the list they are not actually willing to accept.

  • Unstated acceptance is indistinguishable from oversight after an incident
  • Common accepted risks: provider compromise, privileged insiders, credentialed users
  • Naming exclusions during design surfaces the ones nobody actually accepts
  • Scope decisions are business decisions and need a business owner

How to Write an Accepted Risk

An accepted risk entry needs enough structure to be reviewable later. State the scenario concretely, in terms of what an attacker achieves rather than which technique they use. State why you are accepting it — cost of mitigation, low assessed likelihood, the mitigation would break the feature, or the exposure is covered elsewhere. Name the person accepting it, with a date; risk acceptance is an ownership act and an unsigned entry means nobody accepted anything. State what would change the decision: a new data class in the corpus, a new tool with write access, a change in user population, an incident elsewhere in the industry with the same shape. Finally, state any partial measure that remains in place, because most accepted risks are not undefended — they are detected but not prevented, and recording that distinction is what makes the register usable during an incident.

  • Describe the outcome an attacker achieves, not the technique they use
  • Record the reason, the named accepter, and the date — unsigned is unaccepted
  • State the conditions that would reopen the decision
  • Note partial measures: detected-but-not-prevented is a materially different posture

Detection as the Honest Fallback

For a large share of AI-specific risk, prevention is either impossible or so costly that it removes the feature's value — and prompt injection is the obvious example, since no available technique reliably prevents it. The professional response is not to pretend otherwise but to move the control from prevention to containment and detection, and to say so plainly in the model. Containment limits what a successful attack reaches: narrower tools, tighter scopes, denied egress, gated irreversible actions. Detection means you would find out — anomalous tool sequences, unexpected outbound destinations, canary content appearing where it should not, sharp changes in refusal or error rates. Write both into the register alongside the accepted risk. A model that claims injection is mitigated by a filter is worse than one that says injection is expected, contained to these actions, and detected by these signals.

  • Where prevention is unavailable, state containment and detection as the actual controls
  • Claiming injection is prevented by filtering is the most common false assurance in this field
  • Containment is measured by what a successful attack can reach, not by attempt counts
  • Detection commitments belong in the threat model, not only in the runbook

Keeping the Register Alive

Registers die quietly. Three habits keep this one useful. Review on the triggers you wrote down rather than on a calendar — a new tool, a new data class, a new user population, a materially different model deployment — because trigger-based review catches the changes that actually invalidate decisions, while annual review catches whatever is true in the review month. Re-derive acceptance when the underlying assumption changes: a risk accepted because the agent had read-only tools is void the moment it gains a write tool, and that link should be explicit in the entry. And close the loop from incidents, yours and other organisations', by asking whether the scenario appears in your register and whether the acceptance still holds. A register that has never had an entry reopened is not being read.

  • Trigger-based review beats calendar review for catching invalidated assumptions
  • Link each acceptance to the assumption it depends on, so the link breaks visibly
  • Test the register against real incidents, including other organisations'
  • If nothing is ever reopened, the register has become documentation rather than a control

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.