Tool Poisoning Exploits Against MCP Servers
MCP servers can inject malicious instructions through tool metadata before any tool runs.

Tool poisoning is a structural trust flaw in the Model Context Protocol. When an agent connects to an MCP server, it loads that tool's name, description, and schema straight into its context window, and the model treats that metadata the way it treats its own system prompt: as instruction, not as content to question. Nothing in the protocol marks the difference between a plain description ("add two numbers and return the result") and one carrying embedded directives; the two look identical to the agent. The dangerous part is timing. The attack fires the moment a poisoned server enters context, before any tool gets called, and everything below traces how adversaries exploit that opening in practice.
How fast MCP spread and why the attack surface is now enterprise-scale
Anthropic open-sourced MCP in November 2024. Within about six months, every major coding agent, Cursor, Claude Code, Windsurf, Zed, Cline, had built in support for it. OpenAI adopted the protocol in March 2025, Google DeepMind followed soon after, and by December 2025 the Linux Foundation had taken over stewardship. That handoff matters: MCP stopped being one vendor's experiment and became shared infrastructure that no single company owns anymore.
The exposure compounds at enterprise scale, and many risk assessments underestimate how much. IDC projections cited by Microsoft put active AI agents in enterprises at 28.6 million in 2025, climbing past 2.2 billion by 2030. Each of those agents is a potential MCP client pulling tool descriptions into context, and most will draw from servers the organization running the agent didn't build and can't fully check: third-party registries, community-published servers, an integration a teammate wired up last quarter and forgot about. Call it what it is: a supply chain problem. The agent trusts a capability nobody on the team audited, and the more of those agents an enterprise runs, the more capabilities it's trusting sight unseen.
The four ways adversaries weaponize tool descriptions
OWASP has already given this a formal name. Tool poisoning sits in the MCP Top 10 project as MCP03:2025, grouped with rug pulls and tool shadowing under attacks on the capability supply chain agents depend on. Four patterns account for nearly everything observed so far, and they are not interchangeable; each demands a different check, and treating them as one problem is how teams end up fixing the wrong thing.
Description poisoning is the baseline case, and it's the crudest one, since an adversary simply writes malicious directives straight into a tool's description at registration. The agent, loading that metadata to decide what to do next, treats the poisoned text as ground truth about the tool's function. In one reported case, a benign-looking tool's description quietly instructed the agent to read sensitive local files and send them to an attacker. Because the agent executed the newly added server's instructions without a re-prompt, the payload ran on the very next interaction: no extra click, no second warning.
Rug-pull attacks work on a delay, and they're the harder problem precisely because review happens once and trust lasts forever. MCP servers can update their tool list on the fly, and clients refresh either on a schedule or when they get a notifications/tools/list_changed message, so the pattern runs like this: register something harmless, pass whatever review process exists, then swap in a poisoned schema mid-session. Every audit performed against the original clean version becomes worthless the moment the client refreshes and picks up the new definition, which it treats exactly like the one that passed inspection. CVE-2025-54136, disclosed in July 2025 with a CVSS score of 8.8, confirmed this pattern in a production AI development environment: approval of a tool definition did not survive later server-side changes to that same tool. Anyone who treats a one-time review as a permanent guarantee has misread how the protocol works.
Schema poisoning skips the prose and goes straight after the contract. Instead of tampering with a human-readable description, the attacker alters the types, shapes, and rules that define what a request and response actually mean, so an operation that reads as routine at the schema level can map to something destructive underneath. An agent that trusts the schema will misbehave while sailing through any check that only inspects surface validation, and the contract itself changes without a single line of exploited code.
Tool shadowing and cross-server poisoning skip the legitimate server entirely. A malicious MCP server sitting in the same context as a trusted one uses its own descriptions to redirect the agent's calls toward the legitimate server's functions. One documented case involved a WhatsApp MCP server: a co-present malicious server told the agent to pull the user's message history and forward it to an attacker-controlled number, none of it requiring any compromise of WhatsApp's own infrastructure. The attack lives entirely in the second server's metadata and the agent's willingness to follow instructions regardless of where they come from.
Indirect injection through tool outputs: the attack surface that bypasses input validation entirely
In indirect injection, the malicious text arrives inside content the tool fetches and hands back: a document, an email, a web page, a database record. Five vectors cover most observed cases: README or documentation injection, email content injection, web content injection, database record injection, and manipulation of MCP's sampling feature.
Input validation misses all five, because the payload never passes through a field a user typed into. It shows up as the tool's output, and the agent treats tool output as trusted environmental fact, the same category as a file it was asked to read. Nothing in the context window separates "this is a document I was asked to summarize" from "this document contains instructions I should now follow." That's the same ambiguity behind description poisoning, just moved to a different stage of the pipeline.
The practical consequence is blunt: anyone who can write a README, send an email, or drop a record into a database the agent later queries has a potential injection channel, with zero access to the MCP server required. The threat surface for an MCP deployment isn't bounded by the servers an organization runs. It extends to anything those servers might read, which in most real deployments is nearly everything.
Why tool poisoning is not prompt injection by another name
The two get mixed up constantly, and that mix-up deserves attention first. Prompt injection, OWASP's LLM01, rides in user-supplied text, the one channel most security stacks already scan for suspicious content. Tool poisoning rides in server-supplied metadata the model has been told, implicitly, to treat as trusted system-level instruction. They need different defenses because they arrive through different channels entirely, and pointing the same input scanner at both is a common mistake in how teams try to secure MCP today.
The two also differ in how long the damage lasts. A standard prompt injection touches one session and disappears when the conversation ends, while a poisoned tool definition touches every session that loads it, for as long as that server stays connected, which could be months if nobody's watching the connection.
Proximity is the sharpest distinction. A prompt injection needs the attacker to interact with the user or insert something into the conversation directly. Tool poisoning needs neither; access to the tool definition itself, at registration or later through a rug pull, is enough on its own. OWASP catalogs these as separate threat classes, and that separation should drive how a security team spends its attention.
The right way to think about it borrows from software supply-chain security: the adversary corrupts a capability the agent depends on, the same way a malicious package corrupts the code that imports it. Scanning user inputs and hardening conversation-level guardrails, however carefully built, does little to close this particular gap. The tool-metadata channel needs its own enforcement layer, separate from whatever is already watching the chat window.
What empirical benchmarks reveal about how reliably these attacks succeed
The MCPTox benchmark (arXiv 2508.14925, slated for AAAI 2026) is the first systematic attempt to measure how well agents resist tool poisoning under realistic conditions. It runs against 45 live, real-world MCP servers and 353 authentic tools, generates 1,348 malicious test cases across 10 risk categories, and evaluates 20 prominent LLM agents against them.
The average attack success rate landed at 36.5%, meaning more than one attempt in three got through, and the worst recorded case hit 72.8%, against OpenAI's o1-mini. Here's the finding that should worry anyone betting on "better models" as the fix: more capable models were often more susceptible, because the attack exploits good instruction-following rather than a weakness in reasoning. A model that's better at doing what it's told is, by construction, better at doing what a poisoned tool description tells it to do. That flips the usual security instinct, which assumes newer and stronger models are safer by default, and any deployment plan that assumes an upgrade cycle will quietly fix this problem is planning around a fact that isn't true.
Refusal barely factors in. Across the failure cases, agents rarely declined to act on a poisoned instruction; the best refusal rate recorded, from Claude-3.7-Sonnet, sat below 3%. Perhaps the most telling comparison is what happens when the same attack scenarios run through MCP infrastructure versus without it: success rates climb by 7 to 15 percentage points under MCP. Against the InjecAgent benchmark, success rates rose from a 24–48% range outside MCP conditions to 51.2% inside them. That gap says something specific: MCP itself amplifies the risk, apart from whatever vulnerabilities already sit in the underlying model. Safety training built to refuse harmful user requests offers close to no protection here, because the attack arrives through a channel that training was never built to police.
Applying structured threat modeling to MCP deployments
Wiz Research frames the underlying risk as a "lethal trifecta": untrusted input, access to sensitive data, and the ability to act, all three present at once. A typical MCP deployment satisfies all three conditions routinely, often without anyone designing it that way on purpose. That's the pattern worth checking for first, before any tool-by-tool audit: does this agent read something it didn't write, touch data that matters, and take action without a human confirming each step?
Research applying the STRIDE threat-modeling framework has examined five parts of an MCP deployment: the MCP host and client, the LLM itself, the MCP server, external data stores, and the authorization server. Each carries a distinct threat profile, and treating MCP security as one single problem misses where the actual weak points sit. Cisco's integrated AI security framework takes a similar approach, mapping 14 distinct MCP threat types across four groups with assigned severity levels, giving teams a starting point for deciding what to fix first instead of trying to fix everything at once.
The clearest picture of how these attacks unfold in the wild comes from the first cross-entity MCP security study (arXiv 2510.16558), which finds a two-stage attack surface. At the registry level, weak vetting and weak ownership checks let adversarial or hijacked servers into a host before the agent has read a single tool description. At the post-integration level, once a server is in, attacker-controlled tool metadata shapes how the LLM reasons, and the host runs whatever tool calls result without checking them independently. Code-level bugs aren't required for any of this to work, though they can turn a bad situation worse by amplifying attacker-controlled parameters into full exploitation.
An NSA/CISA advisory from May 2026 (CSI_MCP_SECURITY) recommends something deceptively simple: keep a clear inventory of every deployed MCP agent and tool, tracked with versioning, patch history, and known security concerns attached. Skip this step and every other control on this list is theater, since none of them work if nobody knows which servers are connected in the first place. The question worth asking for every MCP server an agent touches is: who can change its tool definitions, when, and with what notice to anyone else on the team?
Controls that address the tool-metadata channel specifically
Pin tool definitions at the point of integration, and don't skip this one. Store a hash or a snapshot of each tool's description and schema at the moment it's approved, and treat any later divergence as an untrusted change requiring re-review before the agent can load it. This is the direct countermeasure to rug-pull attacks; it removes the silent-refresh path that lets a poisoned schema slip in unnoticed. Of every control here, build this one first, because it's the only one that directly blocks CVE-2025-54136's failure mode rather than just catching it after the fact.
Treat tool descriptions as untrusted input during review, full stop. Before connecting a new server, its metadata deserves the same scrutiny a security team gives third-party code: look for embedded directives, conditional logic, references to other tools, or instructions pointing at external destinations that have no business appearing in a tool description.
Enforce least privilege at the level of the tool call, not just the user account. An agent that can read local files and also make outbound network calls in the same session satisfies the lethal trifecta by itself; splitting those permissions across separate tools or separate sessions closes the gap without rewriting the whole system.
Audit the tool-output channel, since most existing security tooling only watches inputs. Indirect injection lives entirely in what a tool sends back to the agent, so logging and checking tool responses matters as much as logging what the user typed, arguably more.
Track every connected MCP server, with version numbers, patch history, and change logs attached. Per the NSA/CISA guidance, this inventory work is what makes rug-pull detection and post-incident triage possible at all, since without knowing which server was running which version when, there's no way to piece together what changed or why.
Re-check approved workflows every time a tool changes. A control that stops an attack but also breaks something people rely on daily gets disabled the first time it's inconvenient, and then it becomes a checkbox nobody enforces. Running those re-checks in an isolated twin environment lets approved workflows be confirmed before anything touches production. Checking that legitimate workflows still work after a safeguard goes in is what keeps that safeguard alive past its first week.
Put a human in the loop for tool definition changes in production. Letting an automated refresh silently adopt a server-side schema change removes the one point in the chain where a person could actually catch a rug pull before it runs. That approval gate is the control itself.


