The AI Learning Hub Journal
◆ Assembly

Context Is Assembled, Not Accumulated

Context is assembled, not accumulatedbuild each call's context fresh from run state — not whatever the transcript has grown intoACCUMULATIONlimitappended to since the run began,trimmed only when something overflows —what to leave out is never decidedomission decisions, quietly deferredwhere agents drift into troublegoalhistoryscratchpadretrievedbudgetassemble(run state)a pure function — testable, versionedthe exact message list for this callevery section present on purpose, sized and orderedTHE MATURITY BENCHMARKPrint the exact context for step nine of a stored run without executing anything, and say why each section is present
If you cannot print a step's exact context without re-running the agent, it is accumulating rather than being assembled

Make It a Function

In most agents that drift into trouble, the context is whatever has accumulated: a message array that has been appended to since the run began, with occasional emergency trimming when something overflows. The alternative is to treat assembly as an explicit function of run state — given the goal, the history, the scratchpad, the retrieved material and the remaining budget, return the exact message list to send. Written that way it becomes ordinary code: testable in isolation, versioned alongside prompts, and inspectable without running the agent. It also forces the questions that accumulation lets you avoid, principally what to leave out. The practical marker of maturity here is that a developer can print the assembled context for step nine of a stored run without executing anything, and can explain why each section is present.

  • Assembly is a pure function of run state, not a side effect of appending
  • Testable, versioned and inspectable without running the agent
  • It forces the omission decisions that accumulation quietly defers
  • Benchmark: can you print step nine's exact context without re-running the agent?

Allocate the Budget by Section

Give each section of the context a nominal token allowance and enforce it during assembly: system instructions, tool definitions, retrieved material, conversation history, scratchpad, and current observations. The allowances do not need to be precise, they need to exist, because the alternative is that one section silently eats the others — a verbose tool result crowding out the retrieved documents that would have made it interpretable. With explicit allowances, overflow becomes a local decision made in the section that overflowed, with a strategy suited to that section: summarise history, drop the lowest-ranked retrieved chunk, truncate the tool result with a pointer to the full artefact. Log the realised token share per section per step. The distribution is almost always surprising the first time a team looks at it, and it usually indicts the tool definitions.

  • Nominal allowances per section, enforced at assembly time
  • Overflow becomes a local decision with a section-appropriate strategy
  • Without allowances, one verbose section quietly evicts the others
  • Log realised token share per section — the first look is usually a surprise

Instrument What Was Actually Sent

The prompt in your template is not the prompt the model received, and the gap between them is where a large share of context bugs live: a truncation that cut a constraint, a retrieval that returned nothing and rendered as an empty section, a tool list assembled in a different order than you assumed, a scratchpad that failed to load. Capture the assembled context per step and store enough of it to reconstruct the run. This overlaps with tracing, covered later as an operational concern, but the motivation here is narrower and more immediate: without it, every context change is evaluated by watching aggregate behaviour and guessing. With it, the first question after any strange run — what did it actually see — has a one-command answer, and most strange runs stop being mysterious at that point.

  • The template is not the prompt; the difference is where context bugs hide
  • Empty retrieval sections and dropped scratchpads render silently
  • Store assembled context per step so "what did it see" is answerable immediately
  • Without this, every context change is evaluated by guesswork over aggregates

Try It Yourself

The benchmark in this lesson is a one-command answer to what the model actually saw at a given step. This is the exercise that tells you whether you have it.

◆ Try it yourself

Take one stored run of an agent you own and print the exact assembled context sent at a middle step - step nine if the run got that far - without re-running anything. Attribute every token to one of the six sections and write down the realised share. If you have no agent, do it ahead of time for the system you are designing: write what the assembly function will emit at step nine and set the nominal allowance per section before any of it is built.

Run:        Step:        Model:        Total input tokens:

Section              | Nominal allowance | Realised tokens | Share | Needed to choose the next action?
System instructions  |  |  |  |
Tool definitions     |  |  |  |
Retrieved material   |  |  |  |
Conversation history |  |  |  |
Scratchpad           |  |  |  |
Current observations |  |  |  |
Unaccounted          |  |  |  |

Then answer:
- Which section is largest, and had you expected it to be?
- Which sections rendered empty or truncated, and would you have noticed from the output?
- Would a careful person given only this context have chosen the same next action?
How you'll know it worked
  • You produced the step's exact context from stored state, with no agent execution
  • Every token is attributed to a named section and the unaccounted row is zero
  • You can name the largest section and say whether the model needed it to choose that step's action

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.