The AI Learning Hub Journal

The Lethal Trifecta as a Design Test

Three properties, checked per component, with evidenceanswer three yes/no questions from live configuration — architectural intent consistently understates accessPRIVATE DATAACCESSUNTRUSTEDCONTENTOUTBOUNDCHANNELALL THREE:EXFILTRATIONENGINEdoes it read anything theuser or world should not see?does anything it reads comefrom outside the boundary?can it cause bytes to reachsomewhere an attacker observes?components with all three go to the top of the list — and the fix is structural, not a better promptREMOVE THE DATA LEGsplit the workflow — the browsingstep gets no customer recordsNARROW THE UNTRUSTED LEGreviewed corpus; schema extractioninstead of prose into the contextREMOVE EGRESS — HIGHEST LEVERAGEdefault-deny is enforced byinfrastructure, outside the modelRUN THREE PASSES — LEAK, ACT, MISLEAD — A CLEAN TRIFECTA IS NOT A CLEAN MODELan agent with no egress can still delete, spend and send; a read-only agent can still confidently mislead
Private data, untrusted content and an outbound channel together make an exfiltration engine — remove one leg rather than defend all three.

Three Properties, Applied Per Component

The most useful design heuristic in this field: a component becomes an exfiltration engine when it combines access to private data, exposure to untrusted content, and a channel to communicate outward. The value is that it converts an unbounded worry into a checkable property. Apply it to every process on your diagram, one at a time, and answer three yes-or-no questions with evidence rather than intuition. Does this component read anything the user or the world should not see — customer records, credentials, other tenants' documents, internal policy? Does anything it reads originate outside your trust boundary, including through the laundering paths of the previous lesson? And can it cause bytes to reach somewhere an attacker can observe? Components with all three go to the top of the remediation list, and the fix is to remove a leg rather than to defend the combination.

  • Private data, untrusted content, outbound channel — checked per component, not per system
  • Answer each with evidence from configuration, not from architectural intent
  • All three present means the fix is structural, not a better prompt
  • Removing one leg is cheaper and more durable than defending all three

Breaking a Leg in Practice

Each leg has real mitigations and real costs, and choosing between them is a product decision as much as a security one. Removing private data access means splitting the workflow: an agent that browses external content gets no customer records, and a separate privileged step operates only on content that has already been reduced to structured, validated fields. Removing untrusted content is rarely possible for a feature whose value is reading external material, but it can be narrowed — a reviewed corpus rather than open browsing, or a fetch step that extracts data through a schema instead of passing prose into the context. Removing the outbound channel is usually the highest-leverage move, because it is enforceable in infrastructure: default-deny egress, no arbitrary URL construction, no auto-rendering of remote content, no write access to shared destinations. Pick the leg you can enforce outside the model.

  • Split workflows so privileged steps operate on validated fields, not raw prose
  • Narrow untrusted exposure where you cannot remove it: reviewed corpora, schema extraction
  • Egress removal is usually most enforceable because infrastructure, not the model, holds it
  • Prefer the leg whose control lives outside the model's decision-making

Where the Heuristic Stops

The trifecta is a design test for exfiltration, and treating it as a complete risk model is the mistake to avoid. Serious harm needs none of the three legs in combination. An agent with no private data access and no egress can still delete records, send an incorrect payment instruction, merge broken code, close the wrong tickets, or post a damaging message — destructive action is a separate axis from data leakage. An agent that only reads can still mislead a user into acting on a fabricated conclusion, and a decision-support system that quietly returns wrong eligibility answers may cause more damage than a leak would. So run the trifecta test, then run a second pass asking what the component can destroy, spend, or send, and a third asking what a confidently wrong answer would cost in this domain. Three passes, three different lists.

  • Exfiltration is one harm class; destruction, spending, and messaging are separate
  • Read-only agents still cause harm by producing confidently wrong conclusions
  • Run three passes: leak, act, mislead — each produces different findings
  • A clean trifecta result is not a clean threat model

Try It Yourself

The test only earns its keep when the three answers come from configuration rather than from what the design intended. One component, half an hour, and the outcome is either a leg you can remove or a risk you have written down.

◆ Try it yourself

Take one component of an AI feature your organisation runs or is designing — one process on the diagram, not the whole system — and answer the three questions with evidence. What private data can it reach, according to the resolved effective permissions of the credential behind each tool rather than the role name? What untrusted content reaches it, including content laundered through an earlier summary or case note? And every way bytes it produces can reach somewhere an attacker can observe: network calls, rendered remote resources, writes to shared stores, messages, third-party query strings, logs and error strings. Then name the leg you could actually remove and what removing it would cost the feature. Section 3 of the AI Feature Threat-Model Canvas at /templates/threat-model-canvas.md is this check. If your organisation has no such feature, run it against an AI feature in a product you use, from its published documentation, and mark each answer as evidenced or assumed. Everything here is answerable by reading configuration; if you want to confirm the egress answer by routing a benign marker, do that only on a system you own or are authorised in writing to test, in a non-production environment with synthetic data.

Component: [one process, not the whole feature]

Trifecta check — evidence, not intent
  Private data reachable: [resolved permission set, not role name]
  Untrusted content reaching it: [channel, including content laundered via an earlier step]
  Outbound paths: [network / rendered resource / shared write / message / query string / log / error string]

All three present? [y/n]
  Leg we can remove:
  What removing it costs the feature:
  Owner and date:

Second pass — what this component can destroy, spend, or send:
Third pass — what a confidently wrong answer costs in this domain:
How you'll know it worked
  • Each of the three answers cites a configuration source — a resolved permission set, an ingestion path, an egress rule — not a design intention
  • Your outbound list includes at least one path that is not a direct network call
  • The component ends with either a named removable leg and its cost, or — if it came out clean — a written note of what it can still destroy, spend, or send

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.