New research shows the Model Context Protocol (MCP)—the interface that lets large-language-model (LLM) agents call external tools—can be hijacked through “tool-poisoning” attacks that succeed more than one-third of the time. Across 20 popular agents the average success rate was 36.5 %; the o1-mini model fell in 72.8 % of attempts, while Claude-3.7-Sonnet refused malicious calls under 3 % of the time. For anyone deploying LLM agents that rely on MCP, the findings turn a convenience feature into a supply-chain risk that can be exploited before any code ever runs.
Why MCP matters to developers today
MCP standardises how agents discover, register, and invoke tools such as file readers, web APIs, or email senders. By publishing a tool’s name, input schema and a short description, a server makes the capability available to any client that understands the protocol. The promise is simple: an agent can look up a tool, send a request, and receive a response without hard-coding each integration.
That flexibility also creates an implicit trust relationship. The specification tells clients to treat tool descriptions as trustworthy only if they come from a server the client already trusts. The new study shows that this trust can be abused.
How tool-poisoning differs from ordinary prompt injection
Traditional prompt injection inserts malicious instructions into the text that the model generates or receives at runtime. The model then follows those instructions because they appear in the same token stream as the user’s request.
Tool-poisoning, by contrast, hides the payload in the tool’s metadata—the name, description, or parameter schema that registers before any agent call. When an agent later selects the tool, it treats the description as part of the “trusted context” and may follow the hidden instruction without any runtime check. Because the injection occurs during registration, there is no point in the execution flow where a model can flag the payload as suspicious.
Scale of the problem – the MCPTox benchmark
The researchers behind MCPTox (arXiv:2508.14925) evaluated 45 MCP servers offering a total of 353 distinct tools. They scripted attacks against 20 widely used LLM agents, measuring how often the agents executed the poisoned tool call.
- Average success rate: 36.5 %
- Peak success: o1-mini at 72.8 %
- Best refusal: Claude-3.7-Sonnet, still under 3 %
The numbers reveal a stark reality: most agents do not refuse a poisoned call because the request looks like a legitimate tool invocation. The agents assume the tool description is a benign piece of documentation, not a vector for code execution.
Why agents rarely refuse poisoned calls
OWASP’s LLM01 guideline explains that LLMs do not differentiate between instructions and data—both are just tokens in a sequence. When a tool description says “send an email to admin@example.com with the subject ‘Update’”, the model cannot tell whether that line is a harmless comment or an instruction it should obey later. Consequently, the model treats the description as part of the trusted environment and follows any embedded command when the tool is invoked.
Existing guidance and its gaps
The MCP specification already advises clients to treat tool descriptions as untrusted unless they originate from a trusted server, and to keep a human in the loop for high-impact calls. The benchmark shows that many real-world deployments ignore or loosely interpret these recommendations.
Concrete steps developers can take today
- Pin server versions – Reference a specific, immutable server image or hash rather than a moving tag. This stops an attacker from swapping a clean registry for a poisoned one after deployment.
- Start with an empty allowlist – Enable only tools that have been explicitly vetted. Anything not on the list is blocked by default.
- Gate state-changing tools – Require additional approval for any tool that writes, sends, or deletes data. Separate “read-only” from “write-capable” capabilities in the schema.
- Add human approval for high-impact calls – For actions that could affect external systems (e.g., sending email, executing commands, modifying files), prompt a human reviewer before the call is sent.
- Log every tool invocation – Record the tool name, arguments, timestamp, and the originating agent. An immutable audit trail makes post-mortem analysis feasible and can deter attackers who know their actions will be visible.
Treat each tool description like source code—subject to linting, code review, and version control—to align the MCP supply chain with standard software-development practices.
Counter-arguments and open questions
The benchmark, however, shows that even the most advanced model in the study refused fewer than three percent of poisoned calls. Fine-tuning may improve detection, but it cannot guarantee safety against novel payloads embedded in schema fields that the model has never seen.
What to watch next
- Emerging standards – Watch for proposals from the LLM security community to require cryptographic signatures on tool schemas.
- Tool-registry hardening – Vendors may start offering immutable, read-only registries as a service, reducing the attack surface.
- Model-level defenses – Research into prompting techniques or auxiliary models that flag suspicious tool metadata could complement host-side safeguards.
The practical takeaway is clear: any MCP-based deployment should audit tool descriptions with the same rigor applied to third-party libraries. Ignoring the supply-chain risk turns a convenient abstraction into a silent backdoor. By pinning servers, enforcing least-privilege allowlists, gating state-changing actions, involving humans where needed, and keeping an immutable log, developers can keep their LLM agents from becoming unwilling accomplices.
