The AI Learning Hub Journal

When a Script Beats an Agent

Put the model where the variance ismark each part of the workflow fixed or variable before choosing an architectureCould you draw the flowchart, given a week?yesnoFIXED — WRITE THE SCRIPTthe steps are known, the order is knownfetch the ticket · parse the payloadlook up the account · apply the rulewrite the record · send the notificationa model choosing these steps adds latency,cost and variance, and removes testabilityVARIABLE — EARN THE LOOPthe next action depends on the last resultdebugging an unfamiliar failureresearch across sources you cannot predictstate you can only discover by lookingclassification, extraction and judgement —where a model beats the alternativesNOT REASONS FOR AN AGENTmany known cases is a branch · several tools is a function · an unclear spec is a design problem the agent inheritsTHE COMMON END STATE — A DETERMINISTIC PIPELINE WITH MODEL-SHAPED HOLESfetchparseclassifyMODELapply ruleextractMODELwritenotifyShip the agent to discover the workflow, then harden the dominant paths into code and keep the loop for the tail
If a week of thought would produce the flowchart, write the flowchart — spend the loop only where the next action depends on the last result

Put the Model Where the Variance Is

The useful decomposition is not whether the task needs an agent but which parts of it are genuinely uncertain. Take any candidate workflow and mark each part as fixed or variable. Fetch the ticket, parse the payload, look up the account, apply the rule, write the record and send the notification are fixed — the steps are known, the order is known, and a model choosing them adds latency, cost and variance while removing your ability to test them. Decide which category this complaint falls into, extract the obligations from this contract, or judge whether this log excerpt indicates the same fault are variable, and a model is genuinely better than the alternatives. The resulting system is a deterministic pipeline with model-shaped holes in it. That is the shape most successful production systems converge on, and it is not usually called an agent.

  • Mark each part of the workflow as fixed or variable before choosing an architecture
  • Fixed steps in a model loop cost latency, money and testability for nothing
  • Classification, extraction and judgement are where a model genuinely wins
  • The common end state is a deterministic pipeline with model-shaped holes

What the Loop Actually Buys

An agent loop earns its cost in exactly one situation: when the next action genuinely depends on the result of the previous one in a way you cannot enumerate in advance. Debugging an unfamiliar failure, researching across sources you cannot predict, or operating in an environment whose state you can only discover by looking are real instances of this. Note what is not on that list. Handling many cases is not it, if the cases are known — that is a branch. Needing several tools is not it — that is a function that calls several tools. Being unsure how to write the logic is not it either, that is a design problem the agent will inherit rather than solve. The test worth applying: could you draw the flowchart if you had a week? If yes, spend the week, because the flowchart will be cheaper, faster and testable forever after.

  • The loop pays off only when the next action truly depends on the last result
  • Many known cases is a branch; several tools is a function call
  • An unclear specification is not an agent use case — the agent inherits the confusion
  • If a week of thought would produce the flowchart, the flowchart wins permanently

Migrating an Agent Back Into Code

The most under-used pattern in this field is treating an agent as a discovery mechanism rather than a destination. Ship an agent to learn what the task actually involves, trace every run, then cluster the trajectories. The dominant paths — and there are usually two or three covering the great majority of traffic — become explicit code, and the agent remains as the fallback for the residual cases that do not match. This is strictly better than either extreme: you get the coverage of an agent while paying agent economics only on the genuinely hard tail, and each hardened path becomes something you can test deterministically and improve without evaluating a whole trajectory. It also converts an uncomfortable conversation about reliability into a graph that goes in the right direction over time.

  • Use the agent to discover the workflow, then harden the dominant paths into code
  • Two or three trajectories usually cover most traffic — those are your functions
  • Keep the agent as the fallback for the residual tail, and pay agent cost only there
  • Each hardened path becomes deterministically testable and independently improvable

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.