The AI Learning Hub Journal

Compaction Without Losing the Thread

Compaction is a protocol, not a summarywhen history no longer fits, what you carry forward decides whether the run keeps its threadFREE-TEXT SUMMARY, THEN CONTINUEask a model to summarise the conversationand resume from the paragraph it wroteTHE CHARACTERISTIC FAILUREthe run continues confidently, a constraintor a decision silently gone, and produceswork that contradicts something agreedtwenty steps earlierEXTRACT INTO A SCHEMA INSTEADgoal as currently understooddecisions taken, with reasonsconstraints still in forceidentifiers touched exactlyopen questionsapproaches tried and rejectedcurrent state of the artefactfilling a schema beats writing a good summaryvalidate it, diff it across compactions, show a humanCOMPACT EARLY AND BY SECTION — SMALL, FREQUENT AND CHEAP, NOT ONE BIG OPERATION AT THE LIMITOLD TOOL OUTPUTcompress aggressively —little of value is lostDECISIONS + CONSTRAINTScarry forward verbatim,never through a summariserDERIVED RUN STATEregenerated at eachassembly — not compactedNARRATIVE MIDDLEthe only part that needsand survives compressionTHE TEST ALMOST NOBODY WRITESCompact a stored run at step N, resume it, and check the agent makes the same next decision it made with full historyLosses cluster on material stated once early and relied on late — keep the pre-compaction state; the diff is the diagnosis
Compact into a schema you can validate and diff, carry decisions verbatim, and test that a resumed run makes the same next decision

Compaction Is a Protocol, Not a Summary

When history no longer fits, the naive move is to ask a model to summarise the conversation and continue from the summary. That works until it does not, and the failure is characteristic: the run continues confidently, having discarded a constraint or a decision, and produces work that contradicts something agreed twenty steps earlier. Treat compaction as a structured extraction into a schema instead of a free-text summary. Fields that earn their place: the goal as currently understood, decisions taken with their reasons, constraints still in force, exact identifiers touched, open questions, approaches already tried and rejected, and the current state of the artefact. Filling a schema is a much more reliable operation than writing a good summary, and it gives you something you can validate, diff between compactions, and show a human.

  • Extract into a schema rather than generating free-text prose
  • Goal, decisions and reasons, live constraints, identifiers, open questions, rejected approaches
  • Schema output can be validated, diffed across compactions and shown to a human
  • The classic failure is confident continuation after silently dropping a constraint

Compact Early, Compact by Section

Waiting until the window is nearly full is the worst time to compact, because you are then doing a large lossy transformation under pressure with no room for the compaction call itself. Compact incrementally instead: retire the oldest tool results into the summary as you go, so the operation is small, frequent and cheap. Prefer sectional compaction over wholesale, since the sections have very different loss tolerances — old tool output compresses aggressively with little cost, whereas decisions and constraints should be carried forward verbatim and never passed through a summariser at all. Anything derived from structured run state does not need compacting because it is regenerated at each assembly. What remains to compress is genuinely just the narrative middle, which is exactly the part that compresses well.

  • Compact incrementally as you go, not in one large operation at the limit
  • Sections have different loss tolerances — compress tool output, carry decisions verbatim
  • Structured run state is regenerated, not compacted
  • The narrative middle is the only part that both needs compression and survives it

Test That the Thread Survived

Compaction is the one context operation with a clean test, and almost nobody writes it. Take a stored run, compact at step N, and resume from the compacted state: does the agent make the same next decision it made with the full history? Run it over the runs you already have and you will find the losses, which are rarely random. They cluster on exactly the material that is stated once early and relied on late — the user's original phrasing of an ambiguous requirement, an identifier mentioned in passing, an exception agreed near the start. Once you know your loss profile you can fix it structurally by promoting those items into the schema. Keep the pre-compaction state as well, because a run that goes wrong after compaction is usually diagnosed by comparing what went in with what came out.

  • Resume a stored run from a compacted state and compare the next decision
  • Losses cluster on material stated once early and relied on late
  • Fix the profile structurally by promoting those items into the schema
  • Retain pre-compaction state — the diff is the diagnosis

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.