Supply Chain Risk in Third-Party MCP Tool Integrations
Tool definitions stay mutable after approval, turning runtime trust into a security vulnerability.

Model Context Protocol turns AI agents into clients that can call file systems, databases, and SaaS tools through one shared connector standard. Most teams are still applying npm-era supply chain thinking to this problem, and that's the wrong model. The trust an agent extends to a tool gets decided at runtime, inside the model's reasoning, not at build time where a scanner can catch it. Lock files and signed hashes were built to freeze something still enough to check. MCP never holds still.
How trust moves at runtime, and why that breaks traditional supply chain controls
Software supply chain security, as practiced for the last decade, revolves around artifacts you can freeze and check. Lock files pin versions. Hashes verify a package hasn't changed since it was signed. Dependency scanners run at compile time, before anything ships. All of that assumes the dangerous part of the system is code, and code sits still long enough to be inspected.
MCP doesn't work that way. An agent reads a tool definition and decides, right there in its reasoning loop, what that tool means and what it's allowed to do. The thing being trusted is a block of natural language processed by a model, not a binary a scanner can walk through. The approval chain has a gap built into it: the model reads the full tool definition underneath, descriptions included, and nothing stops that description from carrying instructions a reviewer never examined. Worse, a server can change that definition after the fact without triggering any new approval step.
Call this what it is: a shift from misinformation risk to mal-action risk. A poisoned tool doesn't hand back a wrong answer a human might catch and correct. It fires a privileged action instead: a file read, an email sent, an API call carrying real credentials. The blast radius tracks the agent's authorization, not the size of the compromised component, so one bad tool definition can do anything the agent itself is allowed to do across every connected service. Classical threat models assume a boundary that holds still. An agent's boundary grows every time it negotiates a new tool or writes something new into memory, and controls built for static dependencies were never designed for a dependency that's a set of instructions the model interprets on its own.
Tool description poisoning: the attack that lives in the semantic layer
Tool poisoning hides malicious instructions inside a tool's metadata (its name, its description, its parameter schema) rather than in anything a user types. The model reads the full tool listing when it plans its next move, and it follows whatever's written there, even when the person who installed the tool never saw the hidden text. This attack class was publicly disclosed in April 2025.
Every field in a tool's schema is a place an attacker can write instructions. Descriptions are the most common target, and a frequent technique wraps the payload in tags that look like markup, so the model reads them as legitimate directives instead of suspicious text. Documented cases have used this to pull SSH keys, MCP configuration files holding credentials for other connected servers, and entire chat histories out through what looks like a normal tool call.
One demonstrated case, involving a WhatsApp MCP integration in April 2025, used a hidden instruction buried in a tool description to exfiltrate an entire chat history. The interface the user saw showed nothing out of the ordinary, because the malicious text never rendered anywhere a human would look.
A related variant hides the exfiltration inside parameters that sound routine, things named context or session_id, with instructions telling the model to fill them with system prompts, prior tool outputs, or conversation history. The data leaves through a normal-looking tool call parameter, so output monitoring built to catch weird text in a chat window never sees it. None of this is code, so static analysis, built to hunt for known-bad code patterns, finds nothing. The attack lives entirely in what the words mean to the model, not in how the file is structured.
Rug-pull attacks: when a tool you already approved turns malicious
A rug-pull starts with a tool that behaves exactly as advertised. It earns adoption, gets approved, builds a track record. Then, once enough trust has piled up, the operator changes the description to add instructions that were never there at review time.
MCP has no built-in mechanism forcing a re-check once a tool has been approved. An operator can edit a description, change what a return value contains, or alter the server's behavior whenever they choose, and the agent keeps treating it as the same integration it trusted on day one. Nothing in the protocol requires a new approval workflow when a tool definition changes after initial review. A poisoned update can go live in production without anyone signing off on it, and without any visible error.
The clearest confirmed case involves a Postmark MCP server, which had around 1,500 weekly downloads, in September 2025. A package impersonating the legitimate server, not an update to it, added a hidden BCC field to its send_email function, quietly copying every outgoing email to an address the attacker controlled. Anyone running it with auto-update turned on leaked email content continuously, with no error message and no change in behavior a user would notice.
This beats a single bad prompt for one reason: persistence changes the math. A rug-pulled tool doesn't compromise one conversation. It compromises every session that touches that server going forward, silently, for as long as it stays connected. Take a routine enterprise pattern: a service owner approves a third-party invoice enrichment tool with no separate security review, the developer later slips a hidden instruction into the tool's description, the update goes live without triggering re-approval, and a financial analyst's ordinary query ends up pulling data it was never supposed to touch. Security review at onboarding checks a moment in time, but the dependency it reviews keeps changing after that moment passes. That mismatch isn't a process failure. It's a structural one.
Tool shadowing: how one malicious server compromises the tools it never touches
Tool shadowing does something stranger than poisoning a single tool. A malicious MCP server injects instructions that change how the agent behaves toward a completely different, legitimate server connected in the same session, without ever touching that other server's code.
A public demonstration showed a compromised server manipulating an agent's handling of a tool belonging to a separate, trusted server, redirecting the agent's actions to benefit an attacker, without any visible indication to the user. The compromised server needed no email capability of its own. It only had to rewrite the shared reasoning context the agent draws on when it decides how to call the trusted tool. This attack class operates at a layer the MCP protocol specification does not address.
The effect spreads sideways instead of staying contained. One compromised server pollutes the shared context that every connected server draws from, opening the door to multi-agent compromise, gradual behavioral drift across otherwise-clean components, and data exfiltration that never touches the tool a defender is watching. Anyone monitoring individual tool calls for anomalies won't find the source, because the legitimate tool executes exactly as designed. It's just following instructions planted somewhere upstream, in context the tool itself never wrote.
Command injection and SDK-level flaws: the vulnerabilities baked into the infrastructure itself
CVE-2025-6514, found in mcp-remote, a package with more than 437,000 downloads, lets a malicious authorization endpoint URL passed through a tool parameter execute arbitrary commands on a shell. The same pattern turned up again in a Figma integration, tracked as CVE-2025-53967. Untrusted input reaching a shell without sanitization is one of the oldest bug classes in software, and here it is again, sitting inside brand-new AI infrastructure.
Research published in April 2026, examining Anthropic's official MCP SDKs across Python, TypeScript, Java, and Rust, found a command execution flaw baked into all four: the STDIO transport passes parameters straight to the host operating system's shell without validating or sanitizing them first. That's not a bug in one company's product. It's a default built into the reference implementation itself, and it propagated into every downstream project that trusted the SDK to handle this correctly. Estimates put the exposure at around 200,000 vulnerable instances, tied to more than 150 million package downloads. Anthropic has confirmed the behavior is intentional and has declined to change the protocol's architecture, leaving individual developers to patch around it on their own. That refusal is the more troubling data point here: the vendor that controls the reference implementation is treating a shell injection path as acceptable by design, and that decision alone should worry anyone deploying MCP servers in production today.
As of May 2026, at least seven confirmed high- or critical-severity CVEs span MCP Inspector, LiteLLM, Cursor, LibreChat, and Windsurf. A scan of the public internet in July 2025 turned up at least 1,862 MCP server instances answering unauthenticated requests, a gap possible because the MCP authorization spec defines OAuth 2.1 but treats authorization as optional rather than required. Given the STDIO flaw, the configuration file becomes the real target: anyone who can edit what an MCP client points to has a path to arbitrary code execution on the host, no exploit chain required beyond that one file.
Registry poisoning, name-squatting, and dependency confusion: the npm mistakes replayed
MCP servers published as npm packages inherit every dependency risk any Node.js package carries, and the Postmark backdoor from September 2025 was the first confirmed case of a malicious MCP server operating in the wild: proof this moved from theoretical concern to something already happening in production.
Name-squatting is documented too. A systematic look at more than 1,800 deployed MCP servers found spoofing and squatting: registering a server whose name closely mirrors a trusted one, aimed at catching agents during configuration or an update cycle. OWASP's Agentic Security Initiative flagged server spoofing as a primary supply chain risk, identifying it as a concern requiring scrutiny beyond individual review.
Dependency confusion is a documented risk here too: many agent frameworks fetch MCP servers through a plain npm install with no check on which registry the package actually came from, leaving them exposed to a well-established class of supply chain attack.
The build pipeline itself is a target too. In June 2025, a path traversal flaw in Smithery, a major hosting platform for MCP servers, exposed builder credentials, including a Docker configuration file containing a Fly.io API token, an exposure that potentially gave attackers reach into more than 3,000 deployed applications. It was publicly disclosed in October 2025. Once the pipeline that builds the servers is compromised, every server that pipeline touches becomes a possible vector, regardless of how careful any individual developer was. None of these patterns are new, and that's the point worth sitting with: what's new is what a compromised package can do once it's running, which is instruct an agent that already has legitimate permission to read files, send email, and call APIs across several enterprise systems at once.
Why the standard threat modeling frameworks don't fully map onto MCP's risk structure yet
Some frameworks do reach into this territory already, and it would be wrong to say nobody's tried. CSA's MAESTRO, published in February 2025, breaks agentic systems into seven layers and names tool interface abuse, supply chain compromise, and privilege escalation as primary threats. Its seventh layer covers the point where ordinary security instincts stop working, because risk emerges from how components combine, not from any one alone. MITRE ATLAS, at version 5.4.0, catalogs 16 tactics, 84 techniques, 56 sub-techniques, 32 mitigations, and 42 case studies, splitting prompt injection into direct and indirect variants under its own technique IDs. OWASP's Top 10 for Agentic Applications, released in December 2025, maps two of its categories, tool misuse and agentic supply chain vulnerabilities, almost directly onto the attack patterns described above, and its earlier Agentic Security Initiative guide from February 2025 laid out a taxonomy spanning agent design, memory, planning, tool use, and deployment.
Even so, the frameworks are behind, and that's the part worth taking seriously rather than glossing over. A systematic review covering 85 papers on agentic LLM security, spanning 2025 and 2026, found research on attacks outpaces research on defenses at a ratio of roughly 3.9 to 1. Risks at the action layer, tool misuse, code injection, sandbox escape, remain understudied relative to how much damage they've already done in the field. No framework yet offers a workable model for continuously re-verifying a tool definition that can change after approval. The attack patterns described above appear in the taxonomies, but the control frameworks meant to stop them haven't caught up. Every framework still assumes something close to a fixed system boundary, while MCP's dynamic tool negotiation means that boundary stretches every time an agent connects to one more server.
Prompt injection sits at the top of OWASP's Top 10 for LLM Applications, and documented research on agentic systems puts its attack success rate as high as 84%. Leading AI developers have publicly noted that the problem is unlikely to ever be fully solved. The gap here isn't spotting the attack. It's what to do once you've spotted it, at the layer where the model turns a plan into an action.
Adoption data backs this up from the governance side. AI tools now show up in the large majority of organizations, by one 2026 industry estimate at 73%, yet real-time enforcement of governance policy reaches only a small fraction of that, around 7% in the same research. The frameworks exist on paper. Putting them into practice at the pace MCP is being adopted is the part nobody has solved.
Controls that actually address what's different about MCP supply chain risk
Treat every MCP server as an untrusted third party from the moment it connects, full stop, and put zero trust to work at the tool integration layer itself, not only at the network perimeter around it. Anything less is a bet that an approved tool stays trustworthy indefinitely, and the Postmark case already proved that bet wrong. That single shift in default posture, recommended in CSA guidance from May 2026, closes off a surprising share of the attack surface described above.
The rug-pull problem needs continuous re-verification, not a one-time review at onboarding. Any change to a tool's name, description, parameter schema, or return value should trigger a fresh approval step before that updated definition reaches a production agent. Pair that with a pinned, version-locked snapshot of each tool definition, a hash or a semantic fingerprint captured at approval time, and you have a concrete baseline to check the live definition against before every session starts, instead of just trusting nothing changed since last time.
Scope matters just as much as verification, and this is where most teams still get it backwards: broad, standing permissions feel efficient right up until one tool definition turns and reaches across every system the agent happens to be authorized to touch. Limit each agent's authorized capabilities to the smallest set its actual workflow requires, and the blast radius of any single compromised tool shrinks down to something containable. Given how the STDIO transport flaw works, avoiding that transport entirely for any server whose configuration file could be altered by an outside party removes the most direct path to arbitrary code execution on the host. None of these controls are exotic. They're the same instincts that built sound infrastructure security for the last twenty years, applied to a dependency type finally forcing the industry to take them seriously again.


