The Trust Curve
Autonomy Is Widened, Not Chosen
The right level of autonomy is not knowable in advance, so treat it as a variable that starts narrow and widens on evidence. A workable progression: the agent proposes and a human executes; then the agent executes with approval on every consequential action; then approval on a defined subset while the rest run free with sampled review; then autonomous operation within limits with review by exception. Each stage produces the evidence needed to justify the next, which is the real reason to run them in order rather than a ceremonial one. Two things make this work in practice. Widen by action type rather than by agent, since an agent can be fully trusted to draft and not at all trusted to send. And write down what would have to be true to move to the next stage before you get there, because the conversation held in advance is a different conversation from the one held under delivery pressure.
- Propose, then approve-all, then approve-subset with sampling, then exception review
- Each stage generates the evidence that justifies the next
- Widen per action type, not per agent — drafting and sending are different questions
- Define the criteria for the next stage before you are standing in front of it
The Evidence That Justifies Widening
Base the decision on something better than an absence of complaints. Useful evidence: the approval rate on the gate you propose to remove, since a gate approved essentially always over a substantial number of decisions is either genuinely safe to relax or is not being read, and you should know which. The disagreement rate in sampled review, which is the direct measure of whether the agent matches human judgement on the cases nobody gated. The failure taxonomy over recent runs, checked specifically for classes with consequences you cannot absorb unattended. And whether the volume reviewed is enough to support any conclusion at all, which is the check most often skipped. Widening also needs a reverse gear: name the conditions that would narrow autonomy again, and make narrowing a routine adjustment rather than an escalation, because a control that can only loosen is not a control.
- Approval rate on the gate in question — and establish whether it was being read
- Disagreement rate from sampled review of ungated runs
- Failure classes present, weighted by whether the consequence is absorbable
- Name the conditions that narrow autonomy again and make narrowing routine
Calibrating What People Expect
Human trust in an automated system moves faster than the evidence in both directions, and both errors are costly. Over-trust builds quietly through a long run of good outcomes and shows up as reviewers who have stopped reviewing while the metrics still say the process is in place. Under-trust follows a single visible failure and produces the opposite waste: an agent that works well being checked exhaustively by people whose attention is needed elsewhere. Counter both with the same practice — publish the actual performance to the people doing the oversight, including the failure rate and the classes of failure, so their calibration is anchored to numbers rather than to the last thing they saw. Telling reviewers what the agent gets wrong, specifically, is also the most effective way to make sampled review productive, because it tells them what to look for.
- Over-trust builds silently through good runs; under-trust follows one visible failure
- Both are miscalibrations against evidence, and both waste scarce attention
- Publish real performance and failure classes to the people doing the oversight
- Telling reviewers what the agent typically gets wrong makes sampling far more productive
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.