Memory by Lifetime
Three Lifetimes, Three Different Problems
Memory is not one feature, it is three stores with genuinely different lifetimes and failure modes, and conflating them is why memory features so often disappoint. The scratchpad lives for one task: notes, the working plan, intermediate results, what has been tried. Its problem is discipline — the agent has to write to it, and it will not unless the loop makes it natural. Task or session memory lives across the interactions that make up one piece of work, and its problem is scoping, since it must be found again by the right run and not by others. Persistent memory lives indefinitely and carries preferences, conventions and accumulated facts, and its problems are curation and staleness. Design each separately, with its own write path, read path and expiry, rather than building one memory system and hoping.
- Scratchpad: one task; the hard part is getting the agent to write to it
- Task or session memory: one piece of work; the hard part is scoping and retrieval
- Persistent memory: indefinite; the hard parts are curation and staleness
- Separate write path, read path and expiry per tier, not one system for all three
Write Policy Is the Hard Part
Reading memory is easy and writing it is where systems go wrong. Left to its own judgement a model writes too much, storing transient details as if they were durable facts, and the store fills with material that is true only of one afternoon. Constrain the write. Give memory a schema with a small number of categories rather than a free-text append. Require a reason and a source for each entry, which both improves quality and makes later review possible. Prefer explicit write moments — end of task, after an approval, on an explicit instruction — over allowing writes at any step. And decide what happens when a new entry contradicts an existing one, because the default of keeping both produces a store that says two things and an agent that picks unpredictably between them.
- Unconstrained writes fill the store with facts that were true for one afternoon
- Schema with limited categories, plus a required reason and source per entry
- Prefer explicit write moments over writes at any arbitrary step
- Define contradiction handling — silently keeping both is how a store starts lying
Retrieval Bounds the Whole Thing
Once memory outgrows what you can paste in wholesale, it is a retrieval system, and its quality ceiling is your retrieval quality rather than anything about the model. That reframing is useful because it imports a body of practice: relevance is measured, recall at the depth you actually inject is the number that matters, and injecting ten marginal memories is worse than injecting two good ones, because the marginal ones consume the same attention budget as the good ones. Two failure modes recur. Memories that surface on topical similarity but not situational relevance, so the agent applies a preference from an unrelated project. And stale entries that outrank fresh ones because they are better written. Timestamp everything, prefer recency where entries conflict, and make the injected memories visible in the trace so you can see which ones influenced a decision.
- Beyond a small store, memory quality is retrieval quality
- Two good memories beat ten marginal ones — attention is the shared budget
- Watch for topical-but-not-situational matches and for well-written stale entries
- Timestamp, prefer recency on conflict, and show injected memories in the trace
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.