The AI Learning Hub Journal

Detecting a Stuck or Looping Agent

Stuck looks exactly like workingfluent reasoning and steady tool calls prove nothing — detection has to be structural, and in the loopREPETITIONthe same tool with materially thesame arguments — it assumes itmade a formatting mistakeOSCILLATIONtwo states alternate and eachaction undoes the last — editand revert, add and removeDRIFTvaried, sensible-looking activitythat never changes the statethe goal is defined byAAAeditrevertSIGNALS YOUR LOOP COMPUTES — ARITHMETIC ON THE TRACE, DURING THE RUNfingerprint = tool name + normalised argumentstwo or three identical fingerprints in a window firesrepeated sequences, not just actions, catch cyclesnear-identical reasoning text means regeneratingprogress predicate: has the success state changed?several flat checkpoints = stuck, however busyWHEN IT FIRES — THREE ESCALATING INTERVENTIONS1 · NAME THE LOOPsay: three identical calls, sameresult — take a different approach2 · REMOVE THE TOOLmake the repeated actionunavailable for the next step3 · STOP AND ESCALATEterminate, and attach the runstate to the escalationWhat buys more looping: an identical retry, a generic "try harder", or a bigger budget
A stuck agent looks exactly like a working one — detect it structurally and escalate the intervention rather than raising the budget

Stuck Looks Exactly Like Working

A stuck agent emits fluent reasoning, calls plausible tools and produces steady log volume. From the outside it is indistinguishable from progress until the budget runs out, which is why detection has to be structural rather than impressionistic. Three shapes cover most of it. Repetition: the same tool with materially the same arguments, over and over, usually because the result is not what the model expected and it assumes it made a formatting mistake. Oscillation: two states alternating, classically edit and revert, where each action undoes the last. And drift without progress: varied, sensible-looking activity that never changes the state you actually care about. Each has a cheap mechanical signature, and none of them require judging the reasoning text.

  • Fluent reasoning and steady tool calls are not evidence of progress
  • Repetition: same tool, same arguments, repeated against an unexpected result
  • Oscillation: alternating actions that undo each other, often edit and revert
  • Drift: plenty of activity, no change in the state that defines the goal

Signals You Can Compute

Fingerprint each action as the tool name plus a normalised form of its arguments, and keep the recent fingerprints in run state. Repeats of the same fingerprint within a short window are the primary signal, and a threshold of two or three is usually right — an agent that reads the same file three times in five steps has stopped learning from the result. Detect cycles by looking for a repeated sequence rather than a repeated single action, which catches the alternating case that per-action counting misses. Add a low-cost content signal: if the model's reasoning text is near-identical to a recent step, it is regenerating rather than reconsidering. All of this is arithmetic on the trace you already keep, and it runs in the loop rather than in a dashboard, because the point is to intervene during the run.

  • Fingerprint = tool name plus normalised arguments; keep a recent window in state
  • Two or three identical fingerprints in a short window is a reliable trigger
  • Look for repeated sequences, not just repeated actions, to catch oscillation
  • Near-identical reasoning text between steps means regeneration, not rethinking

Progress Predicates

The stronger check is task-specific and asks a blunt question at intervals: has the state that defines success changed since the last checkpoint? For a code task, does the diff differ. For a data task, has the count of processed records moved. For a research task, has the set of distinct sources grown. If the answer is no across several consecutive checkpoints, the agent is not working regardless of how much it is doing. This is more work to implement than fingerprinting because it requires you to name what progress means for each task type, and that is precisely why it is valuable: teams that cannot state their progress predicate usually cannot state their completion predicate either, and the two questions are the same question asked at different frequencies.

  • Ask at intervals whether the state defining success has actually changed
  • Diff changed, records processed, distinct sources found — pick a concrete measure
  • Several checkpoints without movement means stuck, however busy the agent looks
  • If you cannot name a progress predicate, you probably cannot name completion either

Intervening Without Making It Worse

Detection is only useful if the response is different from what the agent was already doing. Escalating interventions work well. First, inject an explicit observation into the context — you have called this tool with these arguments three times and received the same result, so treat that result as correct and choose a different approach — because models respond well to being told about a loop they cannot see, given that the repetition is spread across a context they are attending to unevenly. Second, constrain the toolset for the next step so the repeated action is unavailable. Third, terminate and escalate with the state attached. What does not work is retrying identically, adding a generic instruction to try harder, or raising the budget, all of which buy more of the same behaviour at a higher price.

  • Escalate: name the loop in context, then restrict tools, then stop and escalate
  • Models respond well to being told about a repetition they cannot perceive
  • Removing the repeated tool for one step forces a genuinely different path
  • Raising the budget for a looping agent buys more looping, not a result

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.