The AI Learning Hub Journal

What Earns Its Place

What earns its place in the windowsix claimants compete for tokens, and each has to justify itself differentlyTHE PER-STEP TEST — does the model need this to choose the next action?SYSTEM INSTRUCTIONSearn place by changing behaviourparagraphs added after one badoutput have usually never beenshown to change anythingTOOL DEFINITIONSa recurring charge, every stepevery name, description andschema is re-sent each call —an unused tool is pure overheadRETRIEVED MATERIALscope to the step, not the taskfetch what this step needs;re-fetching often beats carryinga document through twenty callsHISTORYaccumulates by defaultthe section most in need of apolicy — raw history is thelowest information densitySCRATCHPADbest value per tokenthe distilled working state —and the most neglected sectionin most agentsCURRENT OBSERVATIONSthe point of the loopwhat just happened is what thenext action turns on — keep itadjacent to generationHISTORY GETS A POLICY, NOT A DEFAULTRecent turns verbatim · older results distilled to outcome and identifiers · failures compressed to the lesson they taught
Every section must justify itself per step against one question — does the model need this to choose the next action

Six Claimants, Different Justifications

Six things compete for the window and each has to justify itself differently. System instructions earn their place if they change behaviour — most long system prompts contain paragraphs that have never been shown to change anything and were added after one bad output. Tool definitions cost tokens on every single step, so an unused tool is a recurring charge. Retrieved material should be scoped to the current step rather than the whole task. History accumulates by default and is the section most in need of a policy. The scratchpad is usually the highest value per token in the whole context and the most neglected. Current observations are the point of the loop. Judge each by a single question asked per step: does the model need this to decide the next action?

  • Judge each section per step: is this needed to choose the next action?
  • Tool definitions are a recurring charge on every step — unused tools are pure overhead
  • Retrieved material should be scoped to the step, not to the whole task
  • The scratchpad is usually the best value per token and the least used

The Tool List Is Context

Teams reason carefully about retrieval and then register thirty tools without a second thought. Every definition — name, description, full argument schema — occupies the window on every step, and the aggregate is frequently larger than the retrieved documents everyone is arguing about. Worse, the cost is not only economic: a long tool list measurably degrades selection, because near-duplicate tools with overlapping descriptions force a discrimination the model has no good basis for making. The remedies are unglamorous. Register only the tools the current phase of the task can use, which is straightforward when the loop has phases. Collapse near-duplicates into one tool with an enumerated mode. And audit usage, because tool lists grow monotonically and nobody has ever removed one without being asked to.

  • Tool definitions often outweigh retrieved material in realised tokens
  • Long lists degrade selection: near-duplicates force an impossible discrimination
  • Register per phase where the loop has phases, rather than everything always
  • Audit call counts — unused tools cost on every step and are never removed spontaneously

History Needs a Policy, Not a Default

Raw history is the least information-dense content in the window: long tool outputs that have already been acted on, reasoning that has been superseded, failed attempts whose only remaining value is one sentence about what did not work. Give it a policy. Keep the last few turns verbatim because recency genuinely matters for coherence. Replace older tool results with a short record of what was fetched and what it established, retaining any identifier the agent may need again. Keep failures, but compressed to the lesson — an agent that loses the memory of a failed approach will retry it, which is one of the more expensive ways to waste a budget. And prefer re-fetching a document over carrying it for twenty steps, since a tool call is usually cheaper than the accumulated cost of transporting the content through every intervening call.

  • Keep recent turns verbatim; distil older tool results to outcome plus identifiers
  • Compress failures to the lesson — a forgotten failure gets retried
  • Re-fetching often costs less than carrying a document through twenty calls
  • Raw history is the lowest information density per token in the whole context

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.