The Honest Limits of Oversight at Volume
At Volume, Humans Cannot Be the Control
Say this plainly, because a lot of AI governance rests on quietly not saying it. Once an agent takes thousands of actions a day, no arrangement of people reviews them meaningfully. What people can do at that scale is review a sample, respond to exceptions, and investigate incidents. What they cannot do is provide assurance on every action, and any design that claims otherwise is describing an aspiration. This is not an argument against human involvement, it is an argument for putting it where it works. Humans set the policy, define what may not happen, review the sample, decide the escalations, and judge whether to widen autonomy. The per-action control at volume has to be structural — deterministic policy checks, scope limits, staged execution with a delay window — because those scale and attention does not.
- At volume, people can sample, handle exceptions and investigate — not assure every action
- A design implying per-action human assurance at scale is describing an aspiration
- Humans set policy, review samples, resolve escalations and decide autonomy
- Per-action control at scale has to be structural, because attention does not scale
Sampling That Discriminates
Uniform random review spends most of its budget confirming that ordinary cases were ordinary. Keep a small random stratum to hold the base rate honest, then spend the rest where the information is: actions at the high end of value or reach, runs where a verification failed or a retry occurred, runs whose step count sits in the tail, cases where the agent expressed low confidence, action types recently promoted to a wider autonomy stage, and anything involving material that arrived from outside. Rotate the emphasis so a stable sampling scheme does not become a predictable blind spot. And close the loop: every review that finds something should produce either a stored case for the test suite or a change to a tool, a schema or a policy, because review that generates no change is a process with no output.
- Keep a small random stratum for the base rate, then sample where information is
- High reach, failed verifications, tail step counts, low confidence, newly widened actions
- Rotate emphasis so the scheme does not become a predictable blind spot
- Every finding should produce a stored case or a change, or review has no output
Oversight Theatre and How to Spot It
The failure state is a process that satisfies an audit and changes nothing, and it has recognisable symptoms. Approval rates at or near total with decision times of a couple of seconds. A sampled review nobody can point to a finding from this quarter. Escalation counts of zero. A policy document naming controls that do not exist in the runtime. An override path used routinely with no record of who used it or why. Each of these is measurable, which is the useful part: oversight quality is not a matter of opinion, it can be instrumented like anything else in this module. Publish those numbers alongside the agent's performance numbers, and treat a gate with a total approval rate as a finding requiring an explanation — either the gate is unnecessary and should be removed, or it is not working and should be fixed. Leaving it in place unexamined is the one option that helps nobody.
- Symptoms: near-total approval, no sampled findings, no escalations, undocumented overrides
- Oversight quality is measurable — instrument it like any other part of the system
- A gate approved essentially always is either unnecessary or not working
- Leaving an unexamined gate in place is the only option with no upside
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.