The AI Learning Hub Journal
◆ Foundations

Why Classic Threat Models Miss AI

Four assumptions that quietly fail when a model sits in the systemthe discipline survives — what changes is the catalogue of instances hanging off each categoryASSUMED: CONTROL FLOW IS FIXEDclassic: enumerate the paths, cover each onethe model chooses the next call at runtime —path enumeration is incomplete by constructionASSUMED: INSTRUCTIONS ≠ DATA, ENFORCEDclassic: the interface separates them, like a parameterised querya prompt is one undifferentiated string —instructions and data share the channelASSUMED: SAME INPUT, SAME OUTPUTclassic: a passing test stays passedsampling makes behaviour a distribution —one green run says little about the nextASSUMED: INSIDE THE BOUNDARY = TRUSTEDclassic: validate at the edge, process freely insideyour own indexed corpus can carry instructionsthat redirect the system reading itTHE METHOD SURVIVES — STRIDE TRANSFERS ONCE YOU SWAP THE INSTANCESspoofing → agent identity · tampering → poisoned corpus and memory · EoP → confused deputy and excessive agencyTHE NEW PRIMITIVE — INFLUENCE WITHOUT ACCESSattacker writes a paragraphof instructionsinto a page, ticket, doc orreview your pipeline ingestsyour system reads it onbehalf of a real userno session to revoke,no credential usedenumerate positions, not identities — user · content author · tool supplier · co-tenant · corpus writerA DECISION-MAKER WITH A PERSUASION SURFACE, NOT A PROGRAM WITH AN ATTACK SURFACEwhat the model can reach once persuaded — not the attacker's identity — sets the severity
Classic threat modelling assumes fixed paths, separable instructions, repeatable behaviour and trustable insides — an LLM feature violates all four, and adds an attacker who never authenticates.

The Assumptions That Break

Classic application threat modelling rests on four assumptions that an LLM feature quietly violates. First, that control flow is fixed, so you can enumerate paths — in an agent, the model chooses the next call at runtime. Second, that instructions and data are separable and the separation is enforced by the interface, the way a parameterised query keeps a value from becoming SQL — a prompt is one undifferentiated string. Third, that identical inputs produce identical outputs, so a passing test stays passed — sampling makes behaviour a distribution, not a fact. Fourth, that trust flows outward from a validated boundary, so once content is inside the system it can be processed safely — but a document from your own indexed corpus can carry instructions that redirect the system that reads it. None of this makes the discipline obsolete. It means the thing you are modelling has stopped being a program with an attack surface and become a decision-maker with a persuasion surface.

  • Control flow is model-chosen at runtime, so path enumeration is incomplete by construction
  • No enforced instruction/data separation — the prompt is one channel carrying both
  • Non-determinism means a single passing test proves very little about the next run
  • Internal content is not trusted content: your own corpus is an attacker delivery path

What Still Works, and Maps Cleanly

Do not throw away the method. Assets, entry points, trust boundaries, data-flow diagrams and STRIDE all transfer — what changes is the catalogue of instances you hang off each category. Spoofing becomes agent and tool identity: can you prove which agent instance, acting for which user, made this call? Tampering becomes poisoning of a retrieval corpus, a memory store, or a tool definition. Repudiation becomes the absence of a trace complete enough to reconstruct why the model acted. Information disclosure becomes exfiltration through outputs and tool side-effects. Denial of service becomes unbounded consumption, since a loop with no ceiling is a budget attack. Elevation of privilege becomes the confused deputy and excessive agency — the agent already holds the privilege, so the attacker only needs to redirect it. Working through STRIDE with these instances in hand is a productive first hour with any team.

  • Spoofing → agent and tool identity; who really made this call, on whose behalf?
  • Tampering → poisoned corpus, poisoned memory, altered tool definitions
  • Information disclosure → exfiltration via output rendering and tool side-effects
  • Elevation of privilege → confused deputy and excessive agency, not a buffer overflow

The New Primitive: Influence Without Access

The genuinely new element is an attacker who never authenticates to your system at all. They write a paragraph into a web page, a support ticket, a shared document, a repository file, a calendar invite, or a product review — anywhere your pipeline will later ingest it — and wait for your system to read it on behalf of a legitimate user. There is no session to revoke and no credential to rotate, because none was ever used. This forces a change in how you enumerate adversaries. The question is no longer only who can call the API, but who can put text in front of the model, directly or eventually, and what the model can do once persuaded. Enumerate positions rather than identities: the end user, the author of any ingested content, the supplier of any tool or server, another tenant whose data shares an index, and the insider who can write to a corpus without review.

  • The attacker may hold no account, no token, and no network path to you
  • Enumerate every party who can influence content your system will eventually read
  • Positions to model: user, content author, tool supplier, co-tenant, corpus writer
  • Ask what the model can reach once persuaded — capability, not identity, sets severity

Scope One Feature, Not the Strategy

Threat modelling sessions fail most often because the scope was set at the wrong altitude. Modelling "our AI programme" produces a list of anxieties; modelling one feature produces decisions. Pick a single deployed or planned capability with a named owner, a defined user population, a specific data set, and an enumerable tool list — the support assistant, the code review bot, the document summariser — and time-box the exercise. Four artefacts should leave the room: a data-flow diagram annotated with trust boundaries, a prioritised list of techniques you consider in scope, a list of controls with named owners and dates, and an explicit statement of what you have decided not to defend against. Also record the triggers that require a revisit — a new tool, a widened permission, a new content source, a change in who the users are. Anything else is a workshop, not a control.

  • One feature, one owner, one data set, one tool list — and a time box
  • Leave with four artefacts: diagram, technique list, control plan, accepted risks
  • Record revisit triggers: new tools, wider permissions, new ingestion sources, new users
  • A threat model with no dated owners attached is a document, not a decision

Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.