Response When the Failure Is a Model
Containment Actions That Exist Here
Familiar containment moves do not apply cleanly: there is no host to isolate and no process to kill, and the component that misbehaved will behave the same way again given the same context. The actions available are about capability and input. Revoke or narrow the agent's tools, which is the fastest way to stop ongoing damage and should be possible without a deployment. Revoke the credentials the agent holds, remembering that a compromised run may have used them for anything within scope. Disable the feature or fall back to a constrained mode. Cut egress to any destination involved. Quarantine the suspected content source — pull the document, pause the sync, freeze the corpus — so the trigger stops arriving. And suspend memory writes and reads for affected scopes, since a persisted instruction will otherwise keep re-entering context after everything else is contained.
- Narrow or revoke tools first — the fastest way to stop ongoing action
- Revoke agent credentials and assume everything in their scope was reachable
- Quarantine the suspected source and cut egress to involved destinations
- Suspend memory reads and writes, or the trigger keeps re-entering context
Scoping Through Traces
Scoping asks which runs were affected and what they did, and it is answerable only from traces. Start from the confirmed case and identify the distinguishing feature — a source document, a retrieved chunk, a tool sequence, a destination — then search all runs for that feature. If a poisoned document is identified, every run that retrieved it is in scope regardless of whether it produced visible harm, and every action taken in those runs must be reviewed. Because indirect injection separates planting from activation, extend the search window back to when the content could have been introduced rather than to when the symptom appeared; scoping to the report date is the standard error here. Then map the affected runs to users and data: whose data was in context, what was returned, what was written, what left the boundary. That mapping is what notification and remediation decisions rest on.
- Find the distinguishing feature, then search all runs for it
- Every run that retrieved the poisoned content is in scope, harm visible or not
- Extend the window back to when the content could have been planted
- Map affected runs to users, data returned, data written, and data that left
Eradication Without a Patch
You cannot patch the model, and that reframes eradication. In most cases the model behaved exactly as it does; what changed was what reached it and what it was permitted to do. So eradication means removing the input and closing the capability. Remove the poisoned content from the corpus and every derived artefact — embeddings, caches, summaries, memory entries — because clearing the source while a summary survives is the failure that makes a cleanup look complete. Purge contaminated memory by origin. Revert prompt or configuration changes if the trigger was introduced there. Then close the capability path: narrow the tool, apply the scope, deny the egress, add the gate. Switching model family or version is occasionally part of the response but rarely the fix on its own, and treating an upgrade as remediation leaves the same path open for the next model.
- The model usually did not change; the input and the permissions are the fix
- Remove derived artefacts as well as sources — embeddings, caches, summaries, memory
- Purge contaminated memory by origin, and verify the purge
- Changing model version is not remediation for a capability path left open
Recovery, Verification, and Communication
Restore the feature in stages: reduced capability first, elevated monitoring, and a defined period before full restoration. Verify with the same benign markers used in testing — re-run the demonstrated path and confirm it now fails, rather than confirming that the specific attempt no longer works, since those are different claims. Communication needs to be prepared for an audience that finds this unfamiliar: explain that the system was manipulated through content it processed rather than breached in the conventional sense, be precise about what data was actually reached, and avoid characterising a systemic property as an isolated defect. Where data was disclosed, obligations follow the data, not the mechanism — regulatory duties do not care that the vector was a prompt. And keep the trace evidence intact and access-controlled, since it contains the same data the incident involved.
- Restore in stages with reduced capability and elevated monitoring
- Verify the path is closed, not that one attempt now fails
- Explain manipulation-through-content plainly; avoid framing systemic issues as isolated defects
- Disclosure obligations follow the data reached, not the novelty of the vector
Try It Yourself
A finding that closes without a test reopens quietly, and a response plan written during an incident is written under pressure. This produces both from one finding.
Take one finding — from your own testing, from an incident in your organisation, or a published class of finding that matches your architecture — and do two things for an AI feature your organisation runs or is designing. First, write it as a regression test that asserts on the trace rather than on output text: that no privileged tool call appears in a run that ingested externally sourced content, that a canary record never reaches a destination, that no request goes to a non-allowlisted host. Make it an invariant asserted across repeated runs, and decide whether it blocks the build or raises a review. Second, write the response steps for the case where the failing component is the model: which tools you would narrow first and whether that is possible without a deployment, which credentials you would revoke, which content source you would quarantine, how far back you would extend the trace search given that planting and activation are separated in time, and which derived artefacts — embeddings, caches, summaries, memory entries — a purge has to clear. Run the test only against a system you own or are authorised in writing to test, in a non-production environment configured like production with synthetic data and dedicated accounts, using benign markers rather than real harm. If your organisation runs no AI feature yet, write both against an AI feature in a product you use and mark every step you could not perform yourself as a dependency on that supplier.
Finding: [what an attacker achieves, and against whom] Condition to test, not the attempt: [the property that must always hold] Regression test Asserts on: [trace field — tools called, acting identity, destination contacted] Runs per change: [n] Tolerance: [zero occurrences / threshold + trend] Environment: [non-production, synthetic data, guardrails as deployed] On failure: [block the build / raise for review] If this fires in production Narrow or revoke first: [tools] Deployment required? [y/n] Credentials to revoke: Source to quarantine: Trace search window: back to [when the content could have been planted] Derived artefacts to purge: embeddings / caches / summaries / memory Verified closed when: [the path fails, not just the original attempt]
- The test asserts on a trace field and would still mean the same thing after a model or prompt change
- The assertion covers repeated runs, and you have decided whether it blocks the build or raises a review
- Your response steps name a containment action that needs no deployment, and a purge that covers derived artefacts as well as the source
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.