The Lethal Trifecta as a Design Test
Three Properties, Applied Per Component
The most useful design heuristic in this field: a component becomes an exfiltration engine when it combines access to private data, exposure to untrusted content, and a channel to communicate outward. The value is that it converts an unbounded worry into a checkable property. Apply it to every process on your diagram, one at a time, and answer three yes-or-no questions with evidence rather than intuition. Does this component read anything the user or the world should not see — customer records, credentials, other tenants' documents, internal policy? Does anything it reads originate outside your trust boundary, including through the laundering paths of the previous lesson? And can it cause bytes to reach somewhere an attacker can observe? Components with all three go to the top of the remediation list, and the fix is to remove a leg rather than to defend the combination.
- Private data, untrusted content, outbound channel — checked per component, not per system
- Answer each with evidence from configuration, not from architectural intent
- All three present means the fix is structural, not a better prompt
- Removing one leg is cheaper and more durable than defending all three
Breaking a Leg in Practice
Each leg has real mitigations and real costs, and choosing between them is a product decision as much as a security one. Removing private data access means splitting the workflow: an agent that browses external content gets no customer records, and a separate privileged step operates only on content that has already been reduced to structured, validated fields. Removing untrusted content is rarely possible for a feature whose value is reading external material, but it can be narrowed — a reviewed corpus rather than open browsing, or a fetch step that extracts data through a schema instead of passing prose into the context. Removing the outbound channel is usually the highest-leverage move, because it is enforceable in infrastructure: default-deny egress, no arbitrary URL construction, no auto-rendering of remote content, no write access to shared destinations. Pick the leg you can enforce outside the model.
- Split workflows so privileged steps operate on validated fields, not raw prose
- Narrow untrusted exposure where you cannot remove it: reviewed corpora, schema extraction
- Egress removal is usually most enforceable because infrastructure, not the model, holds it
- Prefer the leg whose control lives outside the model's decision-making
Where the Heuristic Stops
The trifecta is a design test for exfiltration, and treating it as a complete risk model is the mistake to avoid. Serious harm needs none of the three legs in combination. An agent with no private data access and no egress can still delete records, send an incorrect payment instruction, merge broken code, close the wrong tickets, or post a damaging message — destructive action is a separate axis from data leakage. An agent that only reads can still mislead a user into acting on a fabricated conclusion, and a decision-support system that quietly returns wrong eligibility answers may cause more damage than a leak would. So run the trifecta test, then run a second pass asking what the component can destroy, spend, or send, and a third asking what a confidently wrong answer would cost in this domain. Three passes, three different lists.
- Exfiltration is one harm class; destruction, spending, and messaging are separate
- Read-only agents still cause harm by producing confidently wrong conclusions
- Run three passes: leak, act, mislead — each produces different findings
- A clean trifecta result is not a clean threat model
Try It Yourself
The test only earns its keep when the three answers come from configuration rather than from what the design intended. One component, half an hour, and the outcome is either a leg you can remove or a risk you have written down.
Take one component of an AI feature your organisation runs or is designing — one process on the diagram, not the whole system — and answer the three questions with evidence. What private data can it reach, according to the resolved effective permissions of the credential behind each tool rather than the role name? What untrusted content reaches it, including content laundered through an earlier summary or case note? And every way bytes it produces can reach somewhere an attacker can observe: network calls, rendered remote resources, writes to shared stores, messages, third-party query strings, logs and error strings. Then name the leg you could actually remove and what removing it would cost the feature. Section 3 of the AI Feature Threat-Model Canvas at /templates/threat-model-canvas.md is this check. If your organisation has no such feature, run it against an AI feature in a product you use, from its published documentation, and mark each answer as evidenced or assumed. Everything here is answerable by reading configuration; if you want to confirm the egress answer by routing a benign marker, do that only on a system you own or are authorised in writing to test, in a non-production environment with synthetic data.
Component: [one process, not the whole feature] Trifecta check — evidence, not intent Private data reachable: [resolved permission set, not role name] Untrusted content reaching it: [channel, including content laundered via an earlier step] Outbound paths: [network / rendered resource / shared write / message / query string / log / error string] All three present? [y/n] Leg we can remove: What removing it costs the feature: Owner and date: Second pass — what this component can destroy, spend, or send: Third pass — what a confidently wrong answer costs in this domain:
- Each of the three answers cites a configuration source — a resolved permission set, an ingestion path, an egress rule — not a design intention
- Your outbound list includes at least one path that is not a direct network call
- The component ends with either a named removable leg and its cost, or — if it came out clean — a written note of what it can still destroy, spend, or send
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.