The AI Learning Hub Journal

Verification Inside the Loop

Verification is built into the loop, not requested of the modelan instruction to verify is weakly followed and unauditable — a harness check runs whether the model remembers or notAFTER ANY ACTION THAT CHANGES STATE, THE RUNTIME CHECKSTHE AGENT ACTSan edit, a write, a recordupdate — state has changedA CHEAP DETERMINISTIC CHECKre-read the record, run thetests, parse the file, diffAPPENDED AS AN OBSERVATIONsilent failure becomes acorrection in the same runone extra tool execution, and the run responds to its own failuresSELF-CRITIQUE — WHERE IT HELPSthe check is a different operation from thegenerating — a rubric over an argument, anunhandled case in code, a question that wasonly half-answeredWHERE IT IS CLOSE TO WORTHLESSthe error and the check share the samemisunderstanding — most factual errors, everymisread goal; the reviewer works from thecontext that produced the mistakegive the critique a checklist, not "does this look right" — for higher stakes, run it in a fresh call over the artefact aloneREACH FOR CHECKERS IN THIS ORDER1 · DETERMINISTIC CODEcompiler, linter, validator, tests,checksum — definitive and free2 · JUDGED, WITH A RUBRICwhere the property is realbut not computable3 · A HUMANfor the judgement that isgenuinely irreducibleASK WHAT CODE COULD ESTABLISH BEFORE ASKING WHAT PROMPT COULD ESTABLISHteams skip to the judged middle option and pay model tokens for questions a regular expression would settle
Check work deterministically inside the loop after every state change — completion claims are the weakest evidence available

Build the Check In, Do Not Request It

An instruction to verify your work is weakly followed and impossible to audit. A verification step in the harness runs whether the model remembers or not, and the difference in reliability is large. The pattern that works: after any action that changes state, the runtime performs a cheap deterministic check and appends the result as an observation. Re-read the record and confirm the field. Run the test command. Parse the file. Diff the change. This costs one tool execution and it converts a whole class of silent failure into an observation the agent can respond to in the same run. It also gives you something to assert on later, since a verification result is structured evidence rather than an assertion in prose, which is exactly what the neighbouring practice of evaluation needs to grade a trajectory.

  • A harness check runs regardless of what the model remembers to do
  • After any state change, verify deterministically and append the result as an observation
  • One extra tool execution converts silent failure into an in-run correction
  • Verification results are structured evidence that downstream evaluation can assert on

Self-Critique Has a Narrow Real Range

Asking the model to review its own output does help for some things and is close to worthless for others, and the boundary is worth knowing. It works reasonably where the check is a different cognitive operation from the generation — checking a written argument against a rubric, spotting an unhandled case in code, noticing that an answer failed to address part of the question. It works poorly where the error and the check share the same misunderstanding, which covers most factual errors and every case where the model misread the goal, because the reviewer is working from the same context that produced the mistake. Two practical consequences: give the critique a specific checklist rather than asking whether this looks right, and where the stakes justify the cost, run the critique in a separate call without the working context so it evaluates the artefact rather than the reasoning that produced it.

  • Self-critique helps where checking is a different operation from generating
  • It fails where the error and the check share the same misunderstanding of the goal
  • Give a specific checklist; "does this look right" produces agreement
  • For higher stakes, critique in a fresh call over the artefact alone, without the working context

Prefer the Cheap Deterministic Checker

Before reaching for a model to judge, ask what could be checked by code, because deterministic checks are faster, free, never drift and cannot be argued with. A compiler, a linter, a schema validator, a test suite, a checksum, an arithmetic reconciliation, an existence query — each of these answers a question definitively that a judged check answers probabilistically. The ordering that holds up: deterministic check where one exists, then a judged check with an explicit rubric where the property is real but not computable, then a human where the judgement is genuinely irreducible. Teams often skip straight to the middle option because it is the most flexible, and end up paying model tokens for questions a regular expression would have settled. The habit worth building is asking what code could establish here before asking what prompt could establish here.

  • Deterministic checks are faster, free, stable and unarguable — use them first
  • Then a judged check with an explicit rubric, then a human for the irreducible part
  • Teams default to the judged middle option and pay tokens for regex questions
  • Ask what code could establish before asking what prompt could establish

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.