What Earns Its Place
Six Claimants, Different Justifications
Six things compete for the window and each has to justify itself differently. System instructions earn their place if they change behaviour — most long system prompts contain paragraphs that have never been shown to change anything and were added after one bad output. Tool definitions cost tokens on every single step, so an unused tool is a recurring charge. Retrieved material should be scoped to the current step rather than the whole task. History accumulates by default and is the section most in need of a policy. The scratchpad is usually the highest value per token in the whole context and the most neglected. Current observations are the point of the loop. Judge each by a single question asked per step: does the model need this to decide the next action?
- Judge each section per step: is this needed to choose the next action?
- Tool definitions are a recurring charge on every step — unused tools are pure overhead
- Retrieved material should be scoped to the step, not to the whole task
- The scratchpad is usually the best value per token and the least used
The Tool List Is Context
Teams reason carefully about retrieval and then register thirty tools without a second thought. Every definition — name, description, full argument schema — occupies the window on every step, and the aggregate is frequently larger than the retrieved documents everyone is arguing about. Worse, the cost is not only economic: a long tool list measurably degrades selection, because near-duplicate tools with overlapping descriptions force a discrimination the model has no good basis for making. The remedies are unglamorous. Register only the tools the current phase of the task can use, which is straightforward when the loop has phases. Collapse near-duplicates into one tool with an enumerated mode. And audit usage, because tool lists grow monotonically and nobody has ever removed one without being asked to.
- Tool definitions often outweigh retrieved material in realised tokens
- Long lists degrade selection: near-duplicates force an impossible discrimination
- Register per phase where the loop has phases, rather than everything always
- Audit call counts — unused tools cost on every step and are never removed spontaneously
History Needs a Policy, Not a Default
Raw history is the least information-dense content in the window: long tool outputs that have already been acted on, reasoning that has been superseded, failed attempts whose only remaining value is one sentence about what did not work. Give it a policy. Keep the last few turns verbatim because recency genuinely matters for coherence. Replace older tool results with a short record of what was fetched and what it established, retaining any identifier the agent may need again. Keep failures, but compressed to the lesson — an agent that loses the memory of a failed approach will retry it, which is one of the more expensive ways to waste a budget. And prefer re-fetching a document over carrying it for twenty steps, since a tool call is usually cheaper than the accumulated cost of transporting the content through every intervening call.
- Keep recent turns verbatim; distil older tool results to outcome plus identifiers
- Compress failures to the lesson — a forgotten failure gets retried
- Re-fetching often costs less than carrying a document through twenty calls
- Raw history is the lowest information density per token in the whole context
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.