Verification Inside the Loop
Build the Check In, Do Not Request It
An instruction to verify your work is weakly followed and impossible to audit. A verification step in the harness runs whether the model remembers or not, and the difference in reliability is large. The pattern that works: after any action that changes state, the runtime performs a cheap deterministic check and appends the result as an observation. Re-read the record and confirm the field. Run the test command. Parse the file. Diff the change. This costs one tool execution and it converts a whole class of silent failure into an observation the agent can respond to in the same run. It also gives you something to assert on later, since a verification result is structured evidence rather than an assertion in prose, which is exactly what the neighbouring practice of evaluation needs to grade a trajectory.
- A harness check runs regardless of what the model remembers to do
- After any state change, verify deterministically and append the result as an observation
- One extra tool execution converts silent failure into an in-run correction
- Verification results are structured evidence that downstream evaluation can assert on
Self-Critique Has a Narrow Real Range
Asking the model to review its own output does help for some things and is close to worthless for others, and the boundary is worth knowing. It works reasonably where the check is a different cognitive operation from the generation — checking a written argument against a rubric, spotting an unhandled case in code, noticing that an answer failed to address part of the question. It works poorly where the error and the check share the same misunderstanding, which covers most factual errors and every case where the model misread the goal, because the reviewer is working from the same context that produced the mistake. Two practical consequences: give the critique a specific checklist rather than asking whether this looks right, and where the stakes justify the cost, run the critique in a separate call without the working context so it evaluates the artefact rather than the reasoning that produced it.
- Self-critique helps where checking is a different operation from generating
- It fails where the error and the check share the same misunderstanding of the goal
- Give a specific checklist; "does this look right" produces agreement
- For higher stakes, critique in a fresh call over the artefact alone, without the working context
Prefer the Cheap Deterministic Checker
Before reaching for a model to judge, ask what could be checked by code, because deterministic checks are faster, free, never drift and cannot be argued with. A compiler, a linter, a schema validator, a test suite, a checksum, an arithmetic reconciliation, an existence query — each of these answers a question definitively that a judged check answers probabilistically. The ordering that holds up: deterministic check where one exists, then a judged check with an explicit rubric where the property is real but not computable, then a human where the judgement is genuinely irreducible. Teams often skip straight to the middle option because it is the most flexible, and end up paying model tokens for questions a regular expression would have settled. The habit worth building is asking what code could establish here before asking what prompt could establish here.
- Deterministic checks are faster, free, stable and unarguable — use them first
- Then a judged check with an explicit rubric, then a human for the irreducible part
- Teams default to the judged middle option and pay tokens for regex questions
- Ask what code could establish before asking what prompt could establish
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.