A Human Review Protocol That Is Actually Followed
Why Most Review Protocols Fail
Nearly every firm adopting AI writes a policy requiring human review of AI output. Nearly all of these fail in the same way. The requirement is stated as a general obligation with no named owner, no defined depth, no artefact, and no time budget — and under deadline it degrades into a skim. A protocol that says "all AI output must be reviewed by a qualified lawyer" is not a control, because it does not specify what reviewing means, who does it, how long it takes, or how anyone would know it happened. Controls that survive contact with a busy practice are specific, cheap enough to perform honestly, and leave a trace. Everything else is a policy that exists to be pointed at after something has gone wrong.
- Vague review requirements degrade to skimming under deadline — reliably, not occasionally
- A control needs a named owner, a defined depth, a time budget, and a visible artefact
- If nobody could tell whether the review happened, it is not a control
- A protocol too expensive to follow honestly will be followed dishonestly
The Four Elements That Make It Stick
A workable protocol specifies four things. First, depth by risk tier: an internal summary needs a different check from a document going to a court or a counterparty, and the tiers should be written down with examples. Second, a named individual accountable for each output — not a role, not a team, a person whose name is recorded. Third, an artefact: a verification note, a signed checklist, a comment in the document. The artefact is what makes the control auditable and what makes skipping it a visible act rather than a private one. Fourth, a time allocation in the matter plan, because a review step with no budgeted time is the first thing sacrificed when the deadline moves.
- Depth tiered by risk, with written examples of what each tier requires in practice
- A named person accountable per output — roles and teams diffuse accountability to nobody
- An artefact that records the check, making omission visible rather than private
- Budgeted time in the matter plan, or the step disappears the first time a deadline slips
Sampling What the Tool Did Not Flag
The element most often missing is a check on the tool's coverage rather than its output. Reviewing only what a system flagged tells you about its precision and nothing about its recall, yet recall is where the professional risk lives. A practical protocol includes periodic sampling of unflagged material — a random subset of documents the tool cleared, reviewed blind by a human, with disagreements recorded. Over time this produces something valuable and otherwise unobtainable: an empirical picture of what your tool misses in your document population. That picture is the basis for a defensible statement about method, and it is also how you discover that a model update has quietly changed behaviour on your workload.
- Reviewing only flagged items measures precision; recall is where the professional exposure sits
- Sample unflagged material blind and record disagreements — this is the only way to see misses
- Accumulated disagreement data becomes your defensible evidence about method and limits
- Sampling also detects silent behaviour change after a vendor updates the underlying model
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.