New research shows the Model Context Protocol (MCP)—the interface that lets large-language-model (LLM) agents call external tools—can be hijacked through “tool-poisoning” attacks that succeed more than one-third of the time. Across 20 popular agents the average success rate was 36.5 %; the o1-mini model fell in 72.8 % of attempts, while Claude-3.7-Sonnet refused malicious calls under 3 % of the time. For anyone deploying LLM agents that rely on MCP, the findings turn a convenience feature into a supply-chain risk that can be exploited before any code ever runs.

Why MCP matters to developers today

MCP standardises how agents discover, register, and invoke tools such as file readers, web APIs, or email senders. By publishing a tool’s name, input schema and a short description, a server makes the capability available to any client that understands the protocol. The promise is simple: an agent can look up a tool, send a request, and receive a response without hard-coding each integration.

That flexibility also creates an implicit trust relationship. The specification tells clients to treat tool descriptions as trustworthy only if they come from a server the client already trusts. The new study shows that this trust can be abused.

How tool-poisoning differs from ordinary prompt injection

Traditional prompt injection inserts malicious instructions into the text that the model generates or receives at runtime. The model then follows those instructions because they appear in the same token stream as the user’s request.

Tool-poisoning, by contrast, hides the payload in the tool’s metadata—the name, description, or parameter schema that registers before any agent call. When an agent later selects the tool, it treats the description as part of the “trusted context” and may follow the hidden instruction without any runtime check. Because the injection occurs during registration, there is no point in the execution flow where a model can flag the payload as suspicious.

Scale of the problem – the MCPTox benchmark

The researchers behind MCPTox (arXiv:2508.14925) evaluated 45 MCP servers offering a total of 353 distinct tools. They scripted attacks against 20 widely used LLM agents, measuring how often the agents executed the poisoned tool call.

  • Average success rate: 36.5 %
  • Peak success: o1-mini at 72.8 %
  • Best refusal: Claude-3.7-Sonnet, still under 3 %

The numbers reveal a stark reality: most agents do not refuse a poisoned call because the request looks like a legitimate tool invocation. The agents assume the tool description is a benign piece of documentation, not a vector for code execution.

Why agents rarely refuse poisoned calls

OWASP’s LLM01 guideline explains that LLMs do not differentiate between instructions and data—both are just tokens in a sequence. When a tool description says “send an email to admin@example.com with the subject ‘Update’”, the model cannot tell whether that line is a harmless comment or an instruction it should obey later. Consequently, the model treats the description as part of the trusted environment and follows any embedded command when the tool is invoked.

Existing guidance and its gaps

The MCP specification already advises clients to treat tool descriptions as untrusted unless they originate from a trusted server, and to keep a human in the loop for high-impact calls. The benchmark shows that many real-world deployments ignore or loosely interpret these recommendations.

Concrete steps developers can take today

  1. קיבוע גרסאות שרת – התייחסות לאימג' (image) שרת ספציפי ובלתי ניתן לשינוי או ל-hash, במקום לתג (tag) משתנה. זה מונע מתוקף להחליף רישום (registry) נקי ברישום מורעל לאחר הפריסה.
  2. התחלה עם רשימת הרשאות (allowlist) ריקה – הפעלת כלים שעברו בדיקה מפורשת בלבד. כל דבר שאינו ברשימה נחסם כברירת מחדל.
  3. הצבת מחסומים לכלים המשנים מצב (state-changing tools) – דרישת אישור נוסף עבור כל כלי שכותב, שולח או מוחק נתונים. הפרדה בין יכולות "קריאה בלבד" לבין יכולות "כתיבה" בתוך הסכימה (schema).
  4. הוספת אישור אנושי לקריאות בעלות השפעה גבוהה – עבור פעולות שעלולות להשפיע על מערכות חיצוניות (למשל, שליחת אימייל, הרצת פקודות, שינוי קבצים), יש להציג בקשת אישור למבקר אנושי לפני שליחת הקריאה.
  5. תיעוד (Log) של כל הפעלת כלי – רישום שם הכלי, הארגומנטים, חותמת הזמן והסוכן (agent) המקורי. עקבות ביקורת (audit trail) בלתי ניתנים לשינוי הופכים ניתוח לאחר תקרית (post-mortem) לאפשרי ויכולים להרתיע תוקפים שיודעים שפעולותיהם יהיו גלויות.

יש להתייחס לכל תיאור כלי כאל קוד מקור — הכפוף ל-linting, סקירת קוד ובקרת גרסאות — כדי להתאים את שרשרת האספקה של ה-MCP לפרקטיקות סטנדרטיות של פיתוח תוכנה.

טיעונים נגדיים ושאלות פתוחות

עם זאת, המבחן (benchmark) מראה שגם המודל המתקדם ביותר במחקר סירב לפחות משלושה אחוזים מהקריאות המורעלות. כוונון עדין (fine-tuning) עשוי לשפר את הזיהוי, אך הוא אינו יכול להבטיח בטיחות מפני מטעים (payloads) חדשים המושתלים בשדות סכימה שהמודל מעולם לא ראה.

מה כדאי לעקוב אחריו בהמשך

  • סטנדרטים מתהווים (Emerging standards) – עקבו אחר הצעות מקהילת אבטחת ה-LLM הדורשות חתימות קריפטוגרפיות על סכימות של כלים.
  • הקשחת רישום כלים (Tool-registry hardening) – ספקים עשויים להתחיל להציע רישומים (registries) בלתי ניתנים לשינוי ולקריאה בלבד כשירות, מה שיצמצם את שטח התקיפה.
  • הגנות ברמת המודל (Model-level defenses) – מחקר בטכניקות prompting או מודלים עזר המסמנים מטא-דאטה חשוד של כלים עשוי להשלים את אמצעי ההגנה בצד המארח.

המסקנה המעשית ברורה: כל פריסה מבוססת MCP צריכה לבקר תיאורי כלים באותה קפדנות המופעלת על ספריות צד שלישי. התעלמות מסיכון שרשרת האספקה הופכת הפשטה נוחה לדלת אחורית שקטה. באמצעות קיבוע שרתים, אכיפת רשימות הרשאות של "מינימום הרשאות" (least-privilege), חסימת פעולות המשנות מצב, שיתוף בני אדם במקומות הנדרשים ושמירה על יומן רישום בלתי ניתן לשינוי, מפתחים יכולים למנוע מהסוכנים (agents) של ה-LLM שלהם להפוך לשותפים לא רצויים לעבירה.