The Loop Is Yours, Not the Model's
Who Actually Iterates
The model does not loop. It receives a context, emits one response that may contain tool calls, and stops. Every subsequent iteration happens because your runtime decided to call it again with a larger context. That is not a pedantic distinction, it is the whole design surface: iteration count, what gets appended, what gets dropped, when to stop, and what the model is allowed to do next are all decisions made in your code between calls. Teams that inherit a framework loop without reading it end up debugging behaviour they never chose. Write the loop yourself at least once, even if you later adopt a harness, because the questions it forces you to answer are the questions that determine whether the thing works unattended.
- The model emits one response; your runtime decides whether there is a next one
- Everything between calls — appending, trimming, gating, budgeting — is your code
- A framework loop is still a loop you own the consequences of
- The interesting engineering is in the gaps between model calls, not inside them
What Counts as One Step
Define a step precisely, because budgets, metrics, traces and stuck-detection all key off it. The usable definition: one model call, plus every tool execution that call requested, plus the observations appended before the next call. Under that definition a step can contain several parallel tool calls, which matters because a model that issues four independent reads in one response is doing one step of reasoning with four actions, not four steps. Counting tool executions instead of model calls makes a parallelising agent look wasteful when it is doing the opposite. Pick one definition, write it down, and use it consistently in the step budget, the cost report and the trace, or your numbers will not compare across runs or across features.
- One step = one model call + its tool executions + the resulting observations
- Parallel tool calls inside a response are one step, not several
- Budgets, traces and cost reports must all use the same definition
- Inconsistent step counting makes parallel agents look worse than serial ones
Loop Shapes You Will Actually Choose Between
The interleaved shape — reason, act, observe, repeat — is the default and the most flexible, because each observation can change the plan. Plan-first runs a planning call that produces an explicit ordered list, then executes it with a much smaller model in the loop, which buys predictability, cheaper steps and a plan you can show a human before anything happens, at the cost of adapting badly when reality contradicts the plan. The critique shape inserts a review call over the work so far at fixed points, which catches drift but doubles the cost of every reviewed step. Most production systems are hybrids: plan first, execute interleaved, critique at checkpoints. Choose per feature rather than per company, because the right shape follows from how predictable the task is.
- Interleaved: maximum adaptability, least predictability, the sane default
- Plan-first: an inspectable plan and cheaper execution steps, brittle when the plan is wrong
- Critique at checkpoints: catches drift, and you pay for every review call
- Hybrids are normal — pick the shape from how predictable the task actually is
State Lives Outside the Model
Treat the run as an explicit state object rather than as a conversation you keep appending to. It holds the goal, the remaining budgets, the message history, the scratchpad, the tool results you have chosen to keep, the run identifier, and the status. Keeping it explicit buys three things that matter under unattended operation. It can be persisted, so a run survives a process restart or a wait for human input rather than dying with the request. It can be inspected, so a support engineer can answer what the agent currently believes without reading a transcript. And it can be resumed or forked at a specific step, which is what turns a production failure into something you can actually work on. The conversation is a projection of this state, not the state itself.
- Goal, budgets, history, scratchpad, kept results, run id, status — one object
- Persistable state lets a run pause for a human or survive a restart
- Inspectable state answers "what does it believe now" without reading a transcript
- The message array is a rendering of run state, not the source of truth
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.