Compaction Without Losing the Thread
Compaction Is a Protocol, Not a Summary
When history no longer fits, the naive move is to ask a model to summarise the conversation and continue from the summary. That works until it does not, and the failure is characteristic: the run continues confidently, having discarded a constraint or a decision, and produces work that contradicts something agreed twenty steps earlier. Treat compaction as a structured extraction into a schema instead of a free-text summary. Fields that earn their place: the goal as currently understood, decisions taken with their reasons, constraints still in force, exact identifiers touched, open questions, approaches already tried and rejected, and the current state of the artefact. Filling a schema is a much more reliable operation than writing a good summary, and it gives you something you can validate, diff between compactions, and show a human.
- Extract into a schema rather than generating free-text prose
- Goal, decisions and reasons, live constraints, identifiers, open questions, rejected approaches
- Schema output can be validated, diffed across compactions and shown to a human
- The classic failure is confident continuation after silently dropping a constraint
Compact Early, Compact by Section
Waiting until the window is nearly full is the worst time to compact, because you are then doing a large lossy transformation under pressure with no room for the compaction call itself. Compact incrementally instead: retire the oldest tool results into the summary as you go, so the operation is small, frequent and cheap. Prefer sectional compaction over wholesale, since the sections have very different loss tolerances — old tool output compresses aggressively with little cost, whereas decisions and constraints should be carried forward verbatim and never passed through a summariser at all. Anything derived from structured run state does not need compacting because it is regenerated at each assembly. What remains to compress is genuinely just the narrative middle, which is exactly the part that compresses well.
- Compact incrementally as you go, not in one large operation at the limit
- Sections have different loss tolerances — compress tool output, carry decisions verbatim
- Structured run state is regenerated, not compacted
- The narrative middle is the only part that both needs compression and survives it
Test That the Thread Survived
Compaction is the one context operation with a clean test, and almost nobody writes it. Take a stored run, compact at step N, and resume from the compacted state: does the agent make the same next decision it made with the full history? Run it over the runs you already have and you will find the losses, which are rarely random. They cluster on exactly the material that is stated once early and relied on late — the user's original phrasing of an ambiguous requirement, an identifier mentioned in passing, an exception agreed near the start. Once you know your loss profile you can fix it structurally by promoting those items into the schema. Keep the pre-compaction state as well, because a run that goes wrong after compaction is usually diagnosed by comparing what went in with what came out.
- Resume a stored run from a compacted state and compare the next decision
- Losses cluster on material stated once early and relied on late
- Fix the profile structurally by promoting those items into the schema
- Retain pre-compaction state — the diff is the diagnosis
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.