The AI Learning Hub Journal

Memory and Context Contamination

The attack ends; the influence staysif untrusted content can cause a memory write, one successful injection becomes a standing conditionSESSION 1 — THE EVENThostiledocumentmodelmemory writedoc deletedsession closedtrace goneSESSION 2 … N — THE CONDITIONmemory storeenters context astrusted backgroundsteersactionsthe residue re-enters looking like internal, system-generated contentWHO READS THE STORE SETS THE SEVERITYper-user memorycontained to its ownershared / org-wide storeone user's influence reaches othersmulti-tenant indexa cross-tenant attack pathDESIGN REMOVAL BEFORE THE INCIDENT — TARGETED EXCISION IS THE CAPABILITY YOU NEEDorigin metadatasession, user, source doc,context at write timeexpiry by defaultaged-out entries limitexposure silentlyplain-form inspectionunreadable stores cannotbe auditedtested purge pathsby source, by timewindow, by usera purge that leaves embeddings, caches or summaries behind only appears to workPERSISTENCE TURNS AN EVENT INTO A CONDITION — AND CORPORA ARE MEMORY TOOchunking strips framing, topical shaping controls when content surfaces, and unreviewed indexing is an untrusted write
A memory-scoped injection outlives every trace of the attack and returns looking internal — build origin metadata and full purge before you need them.

Persistence Turns an Event into a Condition

A session-scoped injection ends with the session. A memory-scoped one does not, and that difference is the whole reason to treat memory as a distinct surface. If untrusted content can cause an entry to be written to a store the system reads in future sessions, a single successful attack becomes a standing condition: the hostile document can be deleted, the session closed, the user moved on, and the influence remains. Worse, the residue now arrives with the appearance of internal, system-generated content, since memory is typically injected into context as trusted background rather than as retrieved external material. The practical questions are therefore mechanical. What code paths write to memory? Can the model request a write, and is that request validated? Does an entry record what content was in context when it was created? And can an operator find and remove entries by origin?

  • Session injection ends; memory injection persists after every trace of the attack is gone
  • Memory is usually injected as trusted background, which launders its origin
  • Enumerate write paths and check whether the model can trigger them
  • Entries need origin metadata, or targeted removal is impossible

Retrieval Corpora Are Memory Too

Anything the system retrieves from behaves like memory for these purposes, including the wiki, the ticket system, the shared drive, and any automatically synced third-party source. The write path is the review path: if a document can enter the index without human review, the index is a memory store with untrusted writes. Two specifics deserve attention. Chunking can separate content from its context, so a passage that reads as quoted example text in the original document arrives in the prompt without that framing. And embedding-based retrieval means an attacker can influence when their content surfaces by shaping it toward a target topic, which makes targeting a specific user population feasible without any access to your system. Treat corpus ingestion as a security-relevant pipeline: source allowlisting, provenance metadata per chunk, and the ability to purge by source.

  • If documents can be indexed without review, the corpus is an untrusted-write memory store
  • Chunking strips the framing that made content look like a quotation or an example
  • Attackers can shape content toward a topic to control when it surfaces
  • Require source allowlists, per-chunk provenance, and purge-by-source capability
  • OWASP's LLM Top 10 covers poisoned retrieval under its data poisoning and vector and embedding weaknesses entries

Cross-User and Cross-Tenant Contamination

The severity of contamination depends heavily on who reads the store. Per-user memory limits the effect to its owner and is straightforward to reason about. Shared or organisation-level memory, and any index built from user-contributed content, means one user's influence reaches others — which converts a self-inflicted issue into a cross-user attack and, in a multi-tenant product, a cross-tenant one. The controls follow the scoping decision. Default to the narrowest memory scope that makes the feature work, and require a deliberate decision with a named owner to widen it. Where sharing is genuinely needed, put review between contribution and shared visibility, keep tenant partitions enforced at the storage layer rather than by a filter in the query, and verify partitioning by test rather than by configuration reading, because retrieval layers frequently have paths that bypass the intended filter.

  • Scope drives severity: per-user is contained, shared is a cross-user attack path
  • Widening memory scope should require a deliberate, owned decision
  • Enforce tenant partitioning in storage, not as a filter in the query layer
  • Verify partitioning empirically — bypass paths in retrieval layers are common

Making Contamination Removable

Assume some contamination will occur and design so it can be excised — the capability you need in an incident is targeted removal, and it must exist before the incident. That means every stored entry carries origin metadata: which session, which user, which source document, and which content was in context at write time. It means expiry by default, since an entry that ages out limits exposure without anyone noticing a problem. It means inspection tooling that lets an operator read what the system currently holds for a user in plain form, because a store you cannot read is a store you cannot audit. And it means a tested purge path — by source, by time window, by user — that also clears derived artefacts such as embeddings, caches and summaries. Deleting the source document while a summary of it survives is the failure that makes purges look successful and leaves the influence intact.

  • Origin metadata on every entry: session, user, source, and context at write time
  • Default expiry limits exposure without depending on detection
  • Operators must be able to read current memory in plain form to audit it
  • Purge must clear derived artefacts — embeddings, caches, summaries — or it only appears to work

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.