Tool Poisoning and the MCP Supply Chain
The Tool Definition Is Prompt Content
The mechanism behind tool poisoning is simple and frequently missed: a tool's name, description, parameter names and parameter descriptions are inserted into the model's context so it knows when and how to call the tool. That text is prompt content authored by whoever wrote the server. A description can therefore carry instructions to the model that no human reviewing the tool list in a UI would notice, because interfaces typically show a short label rather than the full schema text the model receives. The consequences follow directly. A hostile description can steer the model to call a tool it should not, to pass different arguments than the user intended, or to alter how it uses other, honest tools in the same session. Review the exact text sent to the model, not the rendered summary — the gap between those two is where this class lives.
- Names, descriptions and parameter docs all enter the context as model-readable text
- UIs show labels; the model sees the full schema — review what the model sees
- A poisoned description can influence the use of other tools in the same session
- Argument steering is as damaging as tool-choice steering and less visible
Supply Chain Properties That Make It Worse
Several properties of the ecosystem amplify the basic mechanism. Tool lists can change during a session, since the protocol supports notifying clients that the list has changed — so what you reviewed at connection time is not necessarily what is in context later. Servers are frequently installed from public sources with the same casual trust once given to small package dependencies, which brings the familiar risks: imitation of a popular name, a maintainer handover, or a server that behaves impeccably until it is widely adopted and then changes. Servers are also transitive: one server may itself call others, so the effective surface is a graph rather than a list. And permissions are commonly granted at server granularity while risk lives at tool granularity, so a server needing one read scope receives a broad grant covering everything it exposes.
- Tool lists can change mid-session — connection-time review is not sufficient
- Name imitation, maintainer change, and post-adoption behaviour change all apply here
- Servers can call other servers; the real surface is a graph you have not drawn
- Grants are usually per-server while risk is per-tool — that mismatch over-privileges by default
- On the standard maps: OWASP's LLM Top 10 supply-chain entry, plus MITRE ATLAS's supply chain compromise techniques
Controls That Apply Classic Discipline
Nothing here requires novel security thinking; it requires applying supply chain discipline to a component people do not yet see as a dependency. Maintain an internal registry of approved servers and block direct installation from arbitrary sources. Pin versions and treat an update as a change requiring review, not as a background refresh. Capture a hash of the full tool schema text at approval time and compare it at every session start, so a changed definition surfaces as an event rather than a surprise — this single control addresses both silent updates and mid-session changes. Review the schema text itself for instruction-like content as part of approval. Grant credentials per tool where the client permits it, and put an egress and policy chokepoint between agents and servers so calls can be inspected, constrained and logged centrally rather than trusted individually.
- Internal registry plus blocked arbitrary installation — the same posture as package management
- Pin versions; treat definition updates as reviewable changes
- Hash the full schema text at approval and compare at every session start
- Route calls through a policy chokepoint so inspection and logging are not per-integration work
Verifying What the Server Actually Does
Approval based on documentation is weak, because the description and the behaviour are authored by the same party. Where the risk justifies it, verify empirically in a controlled environment: run the server with instrumentation and observe what it reads, what it writes, and where it connects, then compare that against its declared purpose. A documentation-search server that opens outbound connections to an unrelated destination, reads files outside its stated scope, or requests credentials it never uses is answering the question for you. Include the negative case: exercise the tool with inputs it should refuse and confirm it enforces its own stated constraints. Record the observed behaviour in the registry entry, because that record is what a future reviewer compares against when the server updates. Behaviour, not description, is the artefact worth storing.
- Documentation and behaviour come from the same author — verify independently
- Instrument a sandboxed run: file reads, writes, and outbound connections
- Compare observed behaviour against declared purpose and credential requests
- Store observed behaviour in the registry as the baseline for future updates
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.