The AI Learning Hub Journal

Errors, Retries and Side Effects

Retry below the loop, and make repeats safetransient noise is handled inside the tool; the model sees only decisions — and every side effect assumes at-least-onceTHE MODEL'S LOOPsees only errors that need a decision — every visibleretry costs a full step and can trigger a strategy changeneeds a decision — surface ittransient — never surfacesTHE TOOL LAYERtimeouts, rate limits, brief unavailability retriedhere with bounded backoff — invisible to the modelRETRIES COMPOSE BADLYan agent that retries a toolwhich retries internally canhammer a struggling dependencyTHE CAP LIVES IN THE RUNTIMEdecide the retry layer pertool, note timing in thedescription, cap totalattempts per tool per runIDEMPOTENCY — AT-LEAST-ONCE IS THE HONEST ASSUMPTION FOR ANY SIDE EFFECTrun id + logical actionstable idempotency keyseen before?no → perform, record the keyyes → return the original resultderive the key from intent, not the argument text the model rendered — and if downstream offers nothing, keep the ledger yourselfFAILURE AFTER PARTIAL COMPLETION — WHERE THE WORST INCIDENTS COME FROMrecord creatednotify failedreported: "failed"model retriestwo recordsprefer atomic operations · else report exactly what did and did not happen, with identifiers · expose the compensating action
Retry transients inside the tool, make every side effect idempotent, and report partial completion exactly rather than as failure

Retry Below the Loop, Not Inside It

Transient infrastructure failures — timeouts, rate limits, brief unavailability — should be handled inside the tool implementation with bounded backoff, and never surfaced to the model. Every retry the model performs costs a full step: a model call, the context, and the risk that it responds to a failure by changing approach rather than repeating. Reserve model-visible errors for conditions requiring a decision. There is a caveat worth respecting: retries at both layers compose badly, and an agent that retries a tool which retries internally can produce a surprising number of attempts against a struggling dependency. Decide the retry layer per tool, document it in the description if it affects timing, and cap total attempts per tool per run in the runtime so no combination of layers can exceed it.

  • Handle transient failures in the tool with bounded backoff, invisibly to the model
  • Every model-visible retry costs a whole step and may trigger a strategy change
  • Surface only errors that genuinely require a decision
  • Cap total attempts per tool per run so layered retries cannot compound

Idempotency Is Not Optional Here

Agents retry, and they retry non-deterministically, which makes at-least-once delivery the honest assumption for any tool with a side effect. That means duplicate sends, duplicate charges, duplicate tickets and duplicate records unless the tool is built to prevent it. The standard mechanism applies: derive a stable idempotency key from the run identifier plus the logical action, have the tool record it, and return the original result on a repeat rather than performing the action again. Make sure the key derives from the intent rather than from the arguments as the model rendered them, since a retry with a reworded free-text field must still be recognised as the same action. Where a downstream system offers no idempotency support, implement the ledger in your tool layer, because leaving this to the model's good judgement will fail at some point.

  • At-least-once is the honest assumption for any tool with a side effect
  • Stable key from run identifier plus logical action; return the original result on repeat
  • Derive keys from intent, not from model-rendered argument text
  • If the downstream system offers nothing, keep the ledger in your tool layer

Partial Effects and the Half-Done Action

The situation that produces the worst incidents is a tool that fails after doing part of its work: the record was created, the notification was not sent, and the error the model receives says the operation failed. The model reasonably retries, and now there are two records. Design against this in the tool. Prefer a single atomic operation where the underlying system allows one. Where it does not, make the tool report exactly what happened rather than a binary failure — created, not notified, here is the identifier — so the next action can complete rather than restart. Where an action genuinely cannot be made atomic or idempotent, make the compensating action available as a tool and say so in the description, because an agent that can undo a partial effect is in a much better position than one whose only option is to try again.

  • The worst incidents come from failures after partial completion, reported as total failure
  • Prefer atomic operations; where impossible, report exactly what did and did not happen
  • Return identifiers from partial success so the next step completes rather than restarts
  • Expose compensating actions as tools where neither atomicity nor idempotency is available

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.