Approval Gates on Irreversible Actions
Gate by Consequence, Not by Category
Approval gates are the control most often implemented badly, and the failure is nearly always in placement. Gating every tool call produces volume, volume produces reflexive approval, and a reflexively approved gate is worse than none because it creates a documented human decision that did not occur. Place gates by consequence instead. The first axis is reversibility: can this be undone, by whom, and how quickly. Sending a message, transferring money, deleting a record, granting a permission, deploying, and publishing are irreversible in the practical sense that a third party has already been affected. The second axis is reach: how many records, whose data, which systems. A short list of genuinely gated actions that people read is worth far more than a comprehensive one they click through, and choosing that short list is a design decision the team should make explicitly and revisit.
- Volume destroys gates; reflexive approval documents a decision nobody made
- Gate on reversibility and reach, not on tool category
- Irreversible means a third party has already been affected
- Fewer, meaningful gates beat comprehensive ones that get clicked through
What the Approver Must See
The approval screen is a security interface, and its content determines whether the decision is real. It must show the concrete action and its exact parameters — the actual recipient, the actual amount, the actual record identifier — rather than a natural-language description of intent, because the description is model output and can be inaccurate whether through error or influence. It should show why the action is proposed, with the specific source content that led to it, so an approver can notice that the instruction originated in an external document. It should show reach: how many records, which systems. And it must have a safe default, which is to do nothing, with a timeout that abandons rather than proceeds. Presenting a persuasive summary without the underlying parameters is how an approver ends up authorising something quite different from what they read.
- Show exact parameters, not a natural-language summary of intent
- Surface the source content that motivated the action, including whether it was external
- State reach: how many records and which systems are affected
- Default to abandon on timeout — never proceed
How Gates Are Defeated
Gates fail in patterned ways, and each pattern has a corresponding test. Splitting: a gated action is decomposed into ungated steps that achieve the same effect, so gate on the effect rather than on the specific tool where you can. Fatigue: raising volume until approvals become automatic, which means approval rate and time-to-decision are security metrics worth monitoring — an approver averaging two seconds is not reading. Misleading rationale: the summary is influenced so the action appears routine, which is why parameters must be shown independently of any generated text. Timing: requesting approval inside a long autonomous run when attention has lapsed. And bypass: an alternative code path that reaches the same effect without the gate, which is the one to test for directly, because gates are usually added at one call site rather than at the capability.
- Splitting: gate the effect, not one call site, and test for alternate paths
- Monitor approval rate and decision time — they measure whether the gate is real
- Show parameters independently of generated rationale to defeat misleading summaries
- Test for ungated code paths reaching the same capability
Alternatives When a Human Cannot Be in the Loop
Human approval does not scale to every deployment, so it is worth knowing the substitutes and their honest limits. Policy checks in the runtime can permit or deny an action deterministically based on parameters, provenance, and context — an amount ceiling, a recipient allowlist, a rule that no privileged action follows untrusted content in the same turn. These are strong because they are structural and cheap because they are automatic. Staged execution helps too: perform the action in a reversible form first, such as a draft, a hold, or a scheduled item with a delay window, so a human or a monitor can intervene before it becomes final. Post-hoc sampled review catches drift but not the individual incident, and should be described that way. What does not work is asking a second model to judge whether the action is safe and treating that as a boundary — it is a signal with an error rate, not a control.
- Deterministic policy checks on parameters and provenance scale where humans cannot
- Staged execution — drafts, holds, delay windows — buys intervention time
- Sampled post-hoc review detects drift, not the individual event
- A model judging an action is a signal with an error rate, not a boundary
Try It Yourself
A gate list is quick to write and easy to over-credit. This exercise checks the two things that decide whether a gate is real: what it is placed on, and whether anyone is reading it.
Pull the live tool list for one agent your organisation runs or is designing, including anything a registered server added, and build the table: what each tool can do, whose identity it acts as, whether its effect is reversible and by whom and how quickly, and whether it is gated. Section 4 of the AI Feature Threat-Model Canvas at /templates/threat-model-canvas.md is this table. Then answer three things the table does not. Which ungated actions are irreversible in the sense that a third party has already been affected? For each existing gate, what is the approver actually shown — the exact parameters, or a generated summary of intent? And what have the approval rate and the median decision time been over the last month? If you want to test whether a gated effect can be reached by an alternative code path, do that only against a system you own or are authorised in writing to test, in a non-production environment with synthetic data and dedicated accounts. If you own no such agent, build the table for an AI feature in a product you use from its documentation and its permission prompts, and record every unknown as ungated.
Agent: Tool list source: [live configuration, pulled on DATE] Tool | What it can do | Acts as | Reversible by whom, how fast | Gated? 1. 2. 3. Irreversible and ungated today: Gated but reachable by another code path: For each gate — what the approver sees: [exact parameters / generated summary] Approval rate, last 30 days: Median decision time: Default on timeout: [abandon / proceed]
- Every tool has an acting identity and a reversibility answer, with no cells left blank
- You can name at least one irreversible action that is currently ungated, or produce the list that shows there is none
- For each existing gate you can state what the approver sees and the median decision time, and say whether that time is long enough to have read it
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.