New research shows the Model Context Protocol (MCP)—the interface that lets large-language-model (LLM) agents call external tools—can be hijacked through “tool-poisoning” attacks that succeed more than one-third of the time. Across 20 popular agents the average success rate was 36.5 %; the o1-mini model fell in 72.8 % of attempts, while Claude-3.7-Sonnet refused malicious calls under 3 % of the time. For anyone deploying LLM agents that rely on MCP, the findings turn a convenience feature into a supply-chain risk that can be exploited before any code ever runs.

Why MCP matters to developers today

MCP standardises how agents discover, register, and invoke tools such as file readers, web APIs, or email senders. By publishing a tool’s name, input schema and a short description, a server makes the capability available to any client that understands the protocol. The promise is simple: an agent can look up a tool, send a request, and receive a response without hard-coding each integration.

That flexibility also creates an implicit trust relationship. The specification tells clients to treat tool descriptions as trustworthy only if they come from a server the client already trusts. The new study shows that this trust can be abused.

How tool-poisoning differs from ordinary prompt injection

Traditional prompt injection inserts malicious instructions into the text that the model generates or receives at runtime. The model then follows those instructions because they appear in the same token stream as the user’s request.

Tool-poisoning, by contrast, hides the payload in the tool’s metadata—the name, description, or parameter schema that registers before any agent call. When an agent later selects the tool, it treats the description as part of the “trusted context” and may follow the hidden instruction without any runtime check. Because the injection occurs during registration, there is no point in the execution flow where a model can flag the payload as suspicious.

Scale of the problem – the MCPTox benchmark

The researchers behind MCPTox (arXiv:2508.14925) evaluated 45 MCP servers offering a total of 353 distinct tools. They scripted attacks against 20 widely used LLM agents, measuring how often the agents executed the poisoned tool call.

  • Average success rate: 36.5 %
  • Peak success: o1-mini at 72.8 %
  • Best refusal: Claude-3.7-Sonnet, still under 3 %

The numbers reveal a stark reality: most agents do not refuse a poisoned call because the request looks like a legitimate tool invocation. The agents assume the tool description is a benign piece of documentation, not a vector for code execution.

Why agents rarely refuse poisoned calls

OWASP’s LLM01 guideline explains that LLMs do not differentiate between instructions and data—both are just tokens in a sequence. When a tool description says “send an email to admin@example.com with the subject ‘Update’”, the model cannot tell whether that line is a harmless comment or an instruction it should obey later. Consequently, the model treats the description as part of the trusted environment and follows any embedded command when the tool is invoked.

Existing guidance and its gaps

The MCP specification already advises clients to treat tool descriptions as untrusted unless they originate from a trusted server, and to keep a human in the loop for high-impact calls. The benchmark shows that many real-world deployments ignore or loosely interpret these recommendations.

Concrete steps developers can take today

  1. Фіксуйте версії серверів – посилайтеся на конкретний, незмінний образ сервера або хеш замість мінливого тегу. Це завадить зловмиснику замінити чистий реєстр отруєним після розгортання.
  2. Починайте з порожнього білого списку – дозволяйте лише ті інструменти, які пройшли явну перевірку. Усе, чого немає в списку, блокується за замовчуванням.
  3. Обмежуйте доступ до інструментів, що змінюють стан – вимагайте додаткового схвалення для будь-якого інструмента, який записує, надсилає або видаляє дані. Розділяйте можливості «тільки для читання» та «з можливістю запису» в схемі.
  4. Додайте людське схвалення для критичних викликів – для дій, що можуть вплинути на зовнішні системи (наприклад, надсилання електронної пошти, виконання команд, зміна файлів), запитуйте підтвердження у рецензента перед надсиланням виклику.
  5. Логуйте кожен виклик інструмента – записуйте назву інструмента, аргументи, мітку часу та агента, що ініціював виклик. Незмінний аудиторський слід робить можливим проведення ретроспективного аналізу (post-mortem) і може стримувати зловмисників, які знають, що їхні дії будуть помітні.

Ставтеся до кожного опису інструмента як до вихідного коду — з лінтингом, переглядом коду та контролем версій — щоб привести ланцюг постачання MCP у відповідність до стандартних практик розробки програмного забезпечення.

Контраргументи та відкриті питання

Однак результати тестування показують, що навіть найдосконаліша модель у дослідженні відхилила менше трьох відсотків отруєних викликів. Тонке налаштування (fine-tuning) може покращити виявлення, але воно не може гарантувати безпеку проти нових корисних навантажень, вбудованих у поля схеми, які модель ніколи раніше не бачила.

На що звернути увагу далі

  • Нові стандарти – стежте за пропозиціями спільноти безпеки LLM щодо вимоги криптографічних підписів для схем інструментів.
  • Посилення захисту реєстрів інструментів – постачальники можуть почати пропонувати незмінні реєстри лише для читання як послугу, що зменшить поверхню атаки.
  • Захист на рівні моделі – дослідження методів промптингу або допоміжних моделей, які позначають підозрілі метадані інструментів, можуть доповнити засоби захисту на стороні хоста.

Практичний висновок очевидний: будь-яке розгортання на основі MCP має проходити аудит описів інструментів з такою ж суворістю, як і сторонніх бібліотек. Ігнорування ризиків ланцюга постачання перетворює зручну абстракцію на прихований бекдор. Фіксуючи сервери, впроваджуючи білі списки з принципом найменших привілеїв, обмежуючи дії, що змінюють стан, залучаючи людей там, де це необхідно, і зберігаючи незмінний лог, розробники можуть запобігти перетворенню своїх LLM-агентів на мимовільних спільників.