The AI Learning Hub Journal

Response When the Failure Is a Model

There is no patch to wait for — remove the input, close the capabilitythe model will behave the same way again given the same context; what changed is what reached it and what it could doisolate the host — there is nonenarrow or revoke the agent's toolskill the process — it will recurquarantine the content sourcepatch the binary — no patch existsclose the capability pathCONTAIN· narrow or revoke tools first· revoke agent credentials — all their scope was reachable· disable or fall back to a constrained mode· cut egress to involved destinations· suspend memory reads and writes for affected scopesSCOPE· find the distinguishing feature: doc, chunk, sequence, destination· every run that touched it is in scope, visible harm or not· window back to when content could have been planted — not to the report date· map runs to users, data read, written, and exfiltratedERADICATE· remove the source and every derived artefact: embeddings, caches, summaries, memory· purge contaminated memory by origin; verify the purge· revert prompt or config changes that introduced it· then close the path: narrow, scope, deny, gateRECOVER· staged restore: reduced capability, more monitoring· verify with benign markers: the path fails, not just the original attempt· explain manipulation through content, plainly· obligations follow the data reached, not the vectorCHANGING MODEL FAMILY OR VERSION IS NOT REMEDIATIONoccasionally part of the response, rarely the fix on its own — an upgrade leaves the same capability path open for the next modelCONTAIN CAPABILITY, QUARANTINE INPUT, SCOPE FROM TRACES, VERIFY THE PATH IS CLOSEDa compromised run could use everything in its credentials' scope — assume it did until traces say otherwise
When the failing component is a model there is nothing to patch: contain by narrowing capability, quarantine the content that triggered it, scope from traces back to the planting, and close the capability path.

Containment Actions That Exist Here

Familiar containment moves do not apply cleanly: there is no host to isolate and no process to kill, and the component that misbehaved will behave the same way again given the same context. The actions available are about capability and input. Revoke or narrow the agent's tools, which is the fastest way to stop ongoing damage and should be possible without a deployment. Revoke the credentials the agent holds, remembering that a compromised run may have used them for anything within scope. Disable the feature or fall back to a constrained mode. Cut egress to any destination involved. Quarantine the suspected content source — pull the document, pause the sync, freeze the corpus — so the trigger stops arriving. And suspend memory writes and reads for affected scopes, since a persisted instruction will otherwise keep re-entering context after everything else is contained.

  • Narrow or revoke tools first — the fastest way to stop ongoing action
  • Revoke agent credentials and assume everything in their scope was reachable
  • Quarantine the suspected source and cut egress to involved destinations
  • Suspend memory reads and writes, or the trigger keeps re-entering context

Scoping Through Traces

Scoping asks which runs were affected and what they did, and it is answerable only from traces. Start from the confirmed case and identify the distinguishing feature — a source document, a retrieved chunk, a tool sequence, a destination — then search all runs for that feature. If a poisoned document is identified, every run that retrieved it is in scope regardless of whether it produced visible harm, and every action taken in those runs must be reviewed. Because indirect injection separates planting from activation, extend the search window back to when the content could have been introduced rather than to when the symptom appeared; scoping to the report date is the standard error here. Then map the affected runs to users and data: whose data was in context, what was returned, what was written, what left the boundary. That mapping is what notification and remediation decisions rest on.

  • Find the distinguishing feature, then search all runs for it
  • Every run that retrieved the poisoned content is in scope, harm visible or not
  • Extend the window back to when the content could have been planted
  • Map affected runs to users, data returned, data written, and data that left

Eradication Without a Patch

You cannot patch the model, and that reframes eradication. In most cases the model behaved exactly as it does; what changed was what reached it and what it was permitted to do. So eradication means removing the input and closing the capability. Remove the poisoned content from the corpus and every derived artefact — embeddings, caches, summaries, memory entries — because clearing the source while a summary survives is the failure that makes a cleanup look complete. Purge contaminated memory by origin. Revert prompt or configuration changes if the trigger was introduced there. Then close the capability path: narrow the tool, apply the scope, deny the egress, add the gate. Switching model family or version is occasionally part of the response but rarely the fix on its own, and treating an upgrade as remediation leaves the same path open for the next model.

  • The model usually did not change; the input and the permissions are the fix
  • Remove derived artefacts as well as sources — embeddings, caches, summaries, memory
  • Purge contaminated memory by origin, and verify the purge
  • Changing model version is not remediation for a capability path left open

Recovery, Verification, and Communication

Restore the feature in stages: reduced capability first, elevated monitoring, and a defined period before full restoration. Verify with the same benign markers used in testing — re-run the demonstrated path and confirm it now fails, rather than confirming that the specific attempt no longer works, since those are different claims. Communication needs to be prepared for an audience that finds this unfamiliar: explain that the system was manipulated through content it processed rather than breached in the conventional sense, be precise about what data was actually reached, and avoid characterising a systemic property as an isolated defect. Where data was disclosed, obligations follow the data, not the mechanism — regulatory duties do not care that the vector was a prompt. And keep the trace evidence intact and access-controlled, since it contains the same data the incident involved.

  • Restore in stages with reduced capability and elevated monitoring
  • Verify the path is closed, not that one attempt now fails
  • Explain manipulation-through-content plainly; avoid framing systemic issues as isolated defects
  • Disclosure obligations follow the data reached, not the novelty of the vector

Try It Yourself

A finding that closes without a test reopens quietly, and a response plan written during an incident is written under pressure. This produces both from one finding.

◆ Try it yourself

Take one finding — from your own testing, from an incident in your organisation, or a published class of finding that matches your architecture — and do two things for an AI feature your organisation runs or is designing. First, write it as a regression test that asserts on the trace rather than on output text: that no privileged tool call appears in a run that ingested externally sourced content, that a canary record never reaches a destination, that no request goes to a non-allowlisted host. Make it an invariant asserted across repeated runs, and decide whether it blocks the build or raises a review. Second, write the response steps for the case where the failing component is the model: which tools you would narrow first and whether that is possible without a deployment, which credentials you would revoke, which content source you would quarantine, how far back you would extend the trace search given that planting and activation are separated in time, and which derived artefacts — embeddings, caches, summaries, memory entries — a purge has to clear. Run the test only against a system you own or are authorised in writing to test, in a non-production environment configured like production with synthetic data and dedicated accounts, using benign markers rather than real harm. If your organisation runs no AI feature yet, write both against an AI feature in a product you use and mark every step you could not perform yourself as a dependency on that supplier.

Finding: [what an attacker achieves, and against whom]
Condition to test, not the attempt: [the property that must always hold]

Regression test
  Asserts on: [trace field — tools called, acting identity, destination contacted]
  Runs per change: [n]   Tolerance: [zero occurrences / threshold + trend]
  Environment: [non-production, synthetic data, guardrails as deployed]
  On failure: [block the build / raise for review]

If this fires in production
  Narrow or revoke first: [tools]   Deployment required? [y/n]
  Credentials to revoke:
  Source to quarantine:
  Trace search window: back to [when the content could have been planted]
  Derived artefacts to purge: embeddings / caches / summaries / memory
  Verified closed when: [the path fails, not just the original attempt]
How you'll know it worked
  • The test asserts on a trace field and would still mean the same thing after a model or prompt change
  • The assertion covers repeated runs, and you have decided whether it blocks the build or raises a review
  • Your response steps name a containment action that needs no deployment, and a purge that covers derived artefacts as well as the source

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.