Agent Engineering: Building the Harness
The scaffolding that turns a model into a working system — loops, context, tools, and the failures that only appear in production
Read it in the library →The Loop
The control flow you own rather than the one the model implies: what a step actually is, how a run ends on purpose, the four budgets that keep it bounded, detecting a stuck or oscillating agent, and the frequent case where a deterministic script is simply the better system.
Context Engineering
The window as a budget you spend on purpose: assembling context in code rather than by accretion, deciding what earns its place, ordering for attention, compacting without losing the thread, memory tiers and their lifetimes, and cache-aware layout as a first-order cost lever.
The Harness
Everything around the model that determines whether it can work: tools designed as APIs for a reader who cannot ask questions, descriptions and schemas that steer selection, return values that enable recovery, retry and idempotency for actions with side effects, verification built into the loop, and delegation that adds capability rather than failure modes.
Making It Hold Up
Operating an agent that runs unattended: designing around non-determinism, tracing well enough to reconstruct a failure, the failure taxonomy you will actually meet and what each one signals, cost and latency as requirements rather than reports, degrading gracefully instead of stopping dead, and what belongs in a log.
Keeping Humans In It
Oversight as a designed part of the system: gates people engage with rather than clear, escalation as a first-class outcome, presenting a run so a human can actually judge it, widening autonomy along an evidence-backed trust curve, the honest limits of review at volume, and writing down what the agent is not allowed to do.
Every lesson is free, with no sign-up. Reading happens in the library, where your progress is saved on your device.