The AI Learning Hub Journal

Context Engineering

The context window is a budgetFIXED WINDOW — hard token ceilingSystem prompt + policyTool definitionsschemas cost tokens tooRetrieved documentsoften the biggest sliceand the least curatedConversation historygrows every single turnwhether it helps or notScratchpad / planthe model's own notesall five compete for the same tokensRETRIEVAL — fetch just in timePull the three relevant chunks, not the corpusIndex quality beats window size, every timeCOMPACTION — summarise the tailOld turns collapse into a short running summaryKeeps the thread, hands the tokens backHYGIENE — trim what you loadLoad tools on demand; cap verbose tool outputEvery unused token is attention you paid forA bigger window does not fix a badly packed one — relevance density is what the model actually reads
What you pack into the window beats how big the window is — curation is the engineering work

The Context Window Is a Budget

Context engineering has emerged as the discipline that quietly determines whether agents work: deciding what occupies the model's context window at each step of the loop. Treat the window as a budget, not a bin. Even with very large windows, three pressures make curation mandatory. Cost — input tokens are billed on every iteration of the loop, so a bloated context is a tax paid per step, not once. Latency — bigger prompts are slower prompts. And attention — models demonstrably degrade on long contexts, losing information buried in the middle and getting distracted by stale or irrelevant content; more context routinely produces worse reasoning. Every token should earn its place: does the model need this, for this step? An agent's effectiveness tracks the signal-to-noise ratio of its context far more closely than the raw amount of information it has been handed.

  • Input tokens are re-billed every loop iteration — context bloat compounds per step
  • Long-context degradation is real: burying the key fact in the middle is how it gets ignored
  • Curate for the step, not the task: most task knowledge is irrelevant to the current decision
  • "Just use the big window" trades money and latency for worse attention — a triple loss

Memory Architectures

Agents outlive their context windows, so information must live somewhere with structure. Three tiers cover practice. Scratchpads: working memory for the current task — running notes, plans, intermediate results — externalised to files or state objects so they survive context turnover; writing a plan down and re-reading it also demonstrably keeps long tasks on track. Persistent memory: knowledge across sessions — preferences, decisions, accumulated project facts — with curation as the hard problem: what is worth keeping, when is it stale, what must be forgettable on demand. Retrieval-backed memory: corpora too large for any window, fetched on demand by search — which makes retrieval quality a direct ceiling on agent quality. The unifying move is the same everywhere: store outside the window, load just-in-time, keep only the relevant slice in context. Subagents extend the same principle across agents — a worker spends its own window on a deep dive and returns only the distilled conclusion.

  • Scratchpad: task-scoped working notes — cheap, effective, and the most underused tier
  • Persistent memory: cross-session knowledge — the challenge is curation and staleness, not storage
  • Retrieval-backed: fetch-on-demand corpora — your retrieval stack bounds your agent
  • Subagent isolation is context engineering: burn a separate window, return a summary

Compaction and Summarisation

Long-running agents eventually face the moment the conversation no longer fits, and how you handle it separates agents that finish marathon tasks from agents that forget why they started. Compaction summarises the transcript so far — decisions made, current state, open questions — and replaces the raw history with the summary plus recent turns. Done well it is renewal; done carelessly it is where tasks silently lose the plot, because summarisation is lossy and the loss is invisible until the agent contradicts a discarded constraint. Compact deliberately: preserve decisions and their reasons, unresolved issues, exact identifiers — file paths, names, values — and discard verbatim tool dumps first, since they are the bulkiest and most re-fetchable content. Related tactics: clear stale tool results aggressively once distilled, prefer re-reading a file over carrying it for twenty turns, and anchor durable constraints in the scratchpad so they survive any number of compactions.

  • Compaction is lossy by definition — engineer what survives, don't trust defaults
  • Preserve decisions, rationale, open threads, and exact identifiers; drop raw tool output first
  • Old tool results are the biggest, cheapest cut — distil and clear them promptly
  • Anchor hard constraints in external notes: compaction should never be able to erase the goal

Cache-Aware Layout — and Why This Beats Model Choice

Prompt caching lets providers reuse computation for a prompt prefix that exactly matches a previous request, at a large discount — which turns context layout into a cost-engineering surface. The rule: stable content first, volatile content last. System prompt, tool definitions, and reference material form the cacheable prefix; conversation and fresh tool results append after it. One dynamic token early — a timestamp in the system prompt, tools registered in shifting order — invalidates the cache for everything downstream, multiplying the cost of an agent that hits the same prefix hundreds of times per task. The strategic point runs deeper than billing: teams routinely reach for a bigger model when an agent underperforms, when the actual failure is contextual — the key fact buried, stale results crowding attention, the constraint compacted away. A mid-tier model with disciplined context reliably beats a frontier model drowning in noise. Exhaust context engineering before paying for model upgrades; it is cheaper and it usually was the problem.

  • Layout for caching: stable prefix (system, tools, references) → volatile suffix (turns, results)
  • One early dynamic token breaks the cache for the whole prompt — audit for timestamps and unstable ordering
  • Agent loops re-send the prefix every step: cache discipline compounds dramatically
  • Diagnose context before upgrading models — most "model too weak" reports are context failures
  • Agent Engineering devotes a whole module to this: assembling context in code, ordering for attention, and memory organised by lifetime
◆ See it for yourself
Open the Context Window in the library →

See what fits, what gets dropped, and what that costs.

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.