Cache-Aware Layout and What Actually Moves Quality
Stable Prefix, Volatile Suffix
Providers can reuse computation for a prompt prefix that exactly matches a previous request, at a substantial discount and a real latency saving. In a single-shot application that is a nice optimisation. In an agent loop it is a structural cost lever, because the loop re-sends the same prefix on every step of every run, so a design that preserves the prefix and one that breaks it differ by a large multiple in both cost and time-to-first-token. The rule follows directly: stable content first, in a fixed order, with everything volatile appended after it. System instructions, tool definitions and durable reference material form the prefix. Conversation, observations and freshly retrieved material go after. Design the layout for this from the start, because retrofitting it means moving content that other parts of the system have come to depend on the position of.
- Prefix reuse compounds across every step of every run, unlike in single-shot use
- Stable ordered prefix: instructions, tool definitions, durable references
- Volatile suffix: turns, observations, fresh retrieval
- It affects latency as well as cost, which matters for anything interactive
What Breaks the Prefix, Usually by Accident
Prefix invalidation is almost always accidental and almost always early in the prompt. A timestamp or a session identifier interpolated into the system instructions invalidates everything after it on every call. A tool list assembled by iterating an unordered collection produces a different order per process, so the cache misses across instances while looking fine locally. Per-user preferences injected at the top rather than after the shared prefix fragment the cache across your user base. Retrieved documents placed before the tool definitions push the volatile content into the stable region. None of these are visible in output quality, which is why they persist. Make cache hit rate a monitored metric with an owner, review it after any prompt change, and treat a sudden drop the way you would treat a latency regression.
- A timestamp in the system prompt invalidates everything after it, every call
- Unordered tool registration silently differs per process and misses across instances
- Per-user content at the top fragments the cache across your entire user base
- Monitor cache hit rate and review it after every prompt change
Context Quality Beats Model Choice More Often Than Expected
The recurring finding, reported by enough independent teams to be treated as a default hypothesis, is that agents underperforming for apparent reasoning reasons are usually underperforming for context reasons. The critical fact was present but positioned where attention degrades. The tool description was ambiguous so the wrong tool was chosen and the trajectory was doomed at step two. A constraint was compacted away. Stale results crowded the observation that mattered. The reflex when quality disappoints is to upgrade the model, which is expensive, slow to evaluate and frequently produces a modest improvement that masks the actual defect. Diagnose first, and diagnose specifically: read the assembled context at the step where the run went wrong and ask whether a careful person with only that context would have made the same mistake. Very often the answer is yes.
- Treat context failure as the default hypothesis before reasoning failure
- Read the assembled context at the step things went wrong, not the whole trace
- Ask whether a careful person given only that context would have erred the same way
- Model upgrades often produce a small improvement that hides the real defect
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.