ML Supply Chain Security advanced 8 min read 6 flashcards

MCP Server Trust and Tool Poisoning

Why a tool description is executable content inside the model's trust boundary, how shadowing and rug-pull attacks exploit the absence of any integrity check on that description, and what intake and pinning look like for a dependency whose interface is prose.

A conventional dependency is code you execute. An MCP server is code you execute plus prose the model obeys. The tool description field, which exists so the model can decide when a tool applies, is delivered into the context window with no marker distinguishing description from instruction, and models do not reliably invent that distinction themselves. Whoever controls a server controls text inside the agent's reasoning, and the protocol offers the client no way to verify that the text it sees today is the text it reviewed at install.

Invariant Labs published the first public proof of concept in April 2025, showing a description field that reads as an innocuous helper while directing the model to read a local file and smuggle its contents out through a normal-looking parameter. Simon Willison's contemporaneous writeup framed the structural problem plainly: MCP puts untrusted text on the same footing as the system prompt, and the combination of private data, untrusted content and an exfiltration path is a design defect rather than a bug in any one server (Willison, 2025, Model Context Protocol has prompt injection security problems).

The three attack shapes

Description injection hides directives in the metadata a user never reads. A client renders the tool name in an approval dialog; the payload lives in the description, which the model reads and the human does not.

Shadowing poisons a tool on server A so that it changes the model's behaviour toward a tool on server B. Nothing on B is compromised. The agent's context is a shared namespace, so a single malicious server in a multi-server configuration is a lever on every other server's tools.

Rug pulls exploit the fact that approval is a one-time event. A server behaves for fifteen versions, accumulates trust, and changes its descriptions or its return values afterwards. The client will not notice, because MCP provides no cryptographic attestation of tool definitions and no requirement that a client re-verify them against what was approved. OWASP tracks this as MCP03:2025 Tool Poisoning (OWASP, 2025, MCP Top 10).

How exposed models actually are

MCPTox is the first benchmark to measure this against reality rather than toy servers: 45 live MCP servers and 353 authentic tools, with poisoned metadata injected into the real catalogue (2025, MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers, arXiv:2508.14925). Attack success varies enormously by model and by attack paradigm, from under 1 percent to around 80 percent. The spread is the finding worth carrying: "our model is resistant" is a claim about one model, one prompt style and one attack family, and it does not survive substitution of any of the three.

The first one in the wild

The theoretical timeline closed in September 2025. A package named postmark-mcp on npm impersonated Postmark's email tooling, published fifteen clean versions, and in version 1.0.16 added a single line that BCC'd every outgoing email to an attacker-controlled address (Postmark, 2025, Security Alert; The Hacker News, 2025). It is a textbook rug pull, and it did not need tool poisoning at all: the server simply did something extra with data the user had already authorised it to handle. An agent that is allowed to send mail is allowed to send mail to anyone.

When it breaks

Approval is a snapshot and dependencies are a stream. Reviewing a server once, at install, gives a guarantee about one version. Without pinning the exact version and re-reviewing on change, the review expires the moment the maintainer publishes again, and the client will not tell you it expired.

Consent dialogs show the wrong field. If the approval UI displays a tool's name and arguments but not its full description, the human is approving a label while the model acts on a paragraph. Any intake process that relies on the user "reviewing the tool" has to define which bytes were reviewed.

Least privilege is bounded by the tool's legitimate purpose. Scoping helps enormously against description injection and not at all against postmark-mcp, where the malicious behaviour lies inside the granted capability. Detecting that requires egress inspection or output auditing, not permission scoping.

More servers is superlinearly worse. Shadowing means risk is not the sum over servers but closer to the product of connections: each added server can influence every other server's tools. A ten-server configuration assembled by a developer who trusts each one individually is not a configuration anyone reviewed.

Local versus remote changes the threat, not the trust. A locally-run server removes a network operator from the picture but still executes a third party's code with the developer's own credentials, which is the same exposure as any npm or PyPI package, with an added channel into the model's reasoning.

References and further reading

Every source this page cites, in the order it cites them. All of them open in a new tab.

  1. Willison, 2025, Model Context Protocol has prompt injection security problems simonwillison.net
  2. OWASP, 2025, MCP Top 10 owasp.org
  3. 2025, MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers, arXiv:2508.14925 arxiv.org
  4. Postmark, 2025, Security Alert postmarkapp.com
  5. The Hacker News, 2025 thehackernews.com
Check yourself

6 flashcards for this concept

Click a card to reveal the answer.

Drill the whole track