Writing Down What the Agent May Not Do
The Prohibition List Is a Real Artefact
Most agent projects have a detailed specification of what the system should do and nothing written about what it must never do, which means the boundary exists only in the assumptions of whoever is currently working on it. Write it down: the actions this agent will never take, the data it will never touch, the systems it will never write to, the thresholds it will never exceed without a person, and the categories of decision that are not delegated at all. Keep it specific enough to implement — never send external communications without approval is implementable; be careful with customer data is not. Producing this list is also the fastest way to surface disagreement inside a team, because the exercise regularly reveals that two engineers held incompatible assumptions about what the thing was allowed to do.
- Specify what must never happen, not only what should happen
- Actions, data, systems, thresholds and non-delegated decision types
- Write it at implementable specificity, not as a principle
- The drafting exercise reliably exposes assumptions the team did not know it disagreed on
It Has to Live in the Runtime
A prohibition in the system prompt is a preference expressed to a probabilistic system. The same prohibition implemented as a check the runtime performs before dispatch is a property of the system. Map every line of the list to its enforcement point — an absent tool, a credential scope, a validation rule, a policy check on parameters, a hard budget — and be honest where no enforcement point exists, because the entries you cannot enforce are exactly the ones worth knowing about. Some genuinely cannot be enforced structurally, and for those the honest answer is a detection with a defined response rather than a sentence in a prompt and a hope. The security discipline covers how these controls are built and layered; the engineering obligation here is that the list and the runtime agree, and that the gap between them is written down rather than assumed away.
- A prohibition in a prompt is a preference; in the runtime it is a property
- Map each entry to its enforcement point: missing tool, scope, validation, policy, budget
- Where no enforcement exists, say so and define a detection and a response instead
- The list and the runtime must agree, and any gap must be recorded rather than assumed away
Ownership, Review and Change
The list needs an owner, a review cadence and a change process, or it becomes a founding document that describes a system nobody has run for a year. Review it when the toolset changes, when autonomy widens, when the agent is pointed at a new data source, and after any incident. Treat additions to the toolset as changes to the prohibition list by default, since a new tool is a new capability and the question of whether anything about it should be forbidden deserves an answer at the point of adding it rather than afterwards. Test the important entries the way you would test anything else: a small suite that attempts the prohibited action and asserts it is refused, run in the same pipeline as everything else. A prohibition nobody has verified since it was written is a claim, and claims are what incident reviews are made of.
- Owner, cadence, and a change process, or it describes a system that no longer exists
- Review on toolset changes, autonomy changes, new data sources, and after incidents
- Adding a tool should trigger the question of what about it must be forbidden
- Test the important prohibitions in the pipeline — an unverified prohibition is a claim
Try It Yourself
This list takes an afternoon and is the quickest way to discover that two people on the same system hold different beliefs about what it may do. Write it for one agent, with the enforcement points named.
Write the prohibition list for one agent you own or are designing. If you own none, write it for an agent your organisation has deployed, from what you can observe. Map every line to its enforcement point, and put the entries you cannot enforce in their own section with a detection and a response beside each. Then take every gated action and classify it by reversibility and reach, naming who approves and what artefact they are shown. Give the finished list to one other engineer on the system and record where you disagreed.
AGENT: NEVER - each line specific enough to implement, not a principle - Actions it will never take: - Data it will never read or write: - Systems it will never write to: - Thresholds it will never exceed without a person: - Decisions never delegated at all: ENFORCEMENT Prohibition | Enforcement point (absent tool / credential scope / validation rule / policy check / hard budget) | Tested in the pipeline? NOT ENFORCEABLE STRUCTURALLY Prohibition | Detection | Defined response | Owner GATED ACTIONS Action | Reversible? by whom, in what window | Reach: records, people and systems affected | Approver | Artefact shown (diff, parameters, sources) | Default on timeout Disagreements found with the second reader:
- Every prohibition maps to a named enforcement point, or sits in the not-enforceable section with a detection beside it
- Each gated action names an approver who exists, and an artefact they see rather than a generated summary
- You can name at least one boundary where you and the second reader had assumed different things
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.