Mapping the Attack Surface
Seven Components, Enumerated Separately
A useful surface map splits the feature into seven components and treats each as its own subject. The model, including which provider and family, whether it is hosted or self-run, and what the deployment can observe about it. The prompt layer: system instructions, templates, few-shot content, and anything injected at assembly time. The retrieval corpus: what is indexed, who can write to it, and how documents get in. The tools: every callable action, its arguments, and the credentials behind it. Memory: anything that persists across turns, sessions, or users. Outputs: everywhere model text is rendered, stored, forwarded, or parsed by another system. And humans: the users, reviewers and approvers whose decisions are part of the control flow. Teams that model only the model miss most of their real exposure, because most of it lives in the other six.
- Model, prompts, corpus, tools, memory, outputs, humans — enumerate each separately
- For each, ask who can write to it and who can read from it
- The model is usually the least interesting component in the map
- Anything that persists or renders is a component, even if nobody called it a feature
Follow the Write Path
The fastest way to find AI-specific exposure is to trace, for every component, how content gets in. For the corpus, that means the ingestion pipeline: who can add a document, is there review, are third-party sources synced automatically, can a user upload a file that is indexed for other users. For memory, it means which code path writes an entry, whether the model can request a write, and whether the write records where the content came from. For tools, it means who publishes the tool definition and whether that definition can change without a deployment. For prompts, it means whether any part of the system prompt is assembled from data — a customer name, a ticket subject, a configuration field — because that is a write path into the instruction layer. Write paths you cannot name are write paths you cannot control.
- For each component list the concrete code path that writes to it
- Automatic sync from third-party sources is an unreviewed write path
- Templated system prompts assembled from data give attackers a foothold in the instruction layer
- A tool definition that can change without a deploy is a live write path into the context
Rank by Reach and Reversibility
Not every element of the map deserves equal attention, and the ranking that predicts real incidents uses two axes. Reach: what a given component can touch — how much data, whose data, and how many downstream systems. Reversibility: whether an action taken through it can be undone, and how quickly. A read-only search tool over public documentation has wide reach and total reversibility, and it rarely produces a serious finding. A tool that sends a message, moves money, deletes a record, merges code, or changes a permission is irreversible in the only sense that matters — the effect has already reached a third party by the time you notice. Rank the map by reach multiplied by irreversibility and spend the session on the top of that list. This ranking, written down, is also what you will use later to decide where approval gates and egress controls actually go.
- Two axes that predict incidents: reach of data and reversibility of action
- Irreversible means a third party has already seen or acted on the effect
- Read-only breadth is cheap to defend; narrow write access rarely is
- The ranking becomes the placement plan for gates, sandboxes and egress rules
Keep the Map Current
Attack surface maps decay faster for AI features than for conventional services, because the surface expands through configuration rather than code. A new connector is added, a corpus gains an automatically synced source, an agent is granted a broader scope to unblock a demo, a tool list grows because an MCP server was registered. None of these look like an architecture change in review, and all of them change the model materially. Two practices keep the map honest. First, make tool and permission inventories generated rather than written: derive them from configuration at build time so drift is visible in a diff. Second, attach the map to the change process — adding a tool, a data source, or a permission scope requires a diff to the map and a short review against the existing technique list. The goal is not ceremony, it is that nobody widens the blast radius silently.
- AI surface grows through configuration, which conventional review often misses
- Generate tool and permission inventories from config so drift shows up in a diff
- Require a map update for new tools, new sources, and new scopes
- Temporary scope widening for a demo is the most common permanent change there is
Try It Yourself
Reading a list of components is not the same as producing one for a system you own. This takes about an hour on a single feature, and the value is in the channels that turn out not to be in the design document.
Pick one AI feature your organisation runs or is designing — one owner, one data set, one tool list. Working from configuration and code rather than design documents, list every channel through which content reaches the model: retrieved documents and their metadata, fetched pages, ticket and email text, filenames and document properties, tool responses including error strings, memory entries, and anything templated into the system prompt. Mark each channel trusted or untrusted, name the concrete code path that writes to it, and note whether any test exercises it. Section 2 of the AI Feature Threat-Model Canvas at /templates/threat-model-canvas.md is this table, so start from it rather than building a scaffold. If your organisation ships no AI feature yet, build the same inventory for an AI feature in a product you use, from its published documentation, and record every channel you cannot confirm as unknown rather than trusted. This is a configuration and documentation review only — do not probe any system you are not authorised in writing to test.
Feature: Owner: Inventory built from: [configuration and code, on DATE] Channel | Trusted or untrusted | Code path that writes to it | Exercised by a test? 1. 2. 3. Channels found that were not in the design document: Channels nobody could name a write path for: Untrusted channels no test currently exercises:
- Your inventory contains at least one channel that was not in the feature's design document
- Every channel has a named write path, or is recorded as one nobody could name
- You can point to at least one untrusted channel that no existing test exercises
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.