Est.
Agent ThreatsLong read

OWASP Agentic Top 10 Applied to Production Deployments

Real agents with real credentials demand controls that survive when the model's goals get hijacked.

Staff Writer · · 11 min read · Updated
Cover illustration for “OWASP Agentic Top 10 Applied to Production Deployments”
Agent Threats · September 11, 2026 · 11 min read · 2,402 words

The OWASP Agentic Top 10, published December 9, 2025, names ten failure modes that show up when an AI system stops answering and starts acting. Built with more than 100 industry experts and practitioners, it maps what happens when an agent holds real credentials, calls real tools, and carries a plan across multiple steps instead of a single response. This piece walks through five of those categories with an operational question in mind: what does the risk look like in a live deployment, and what does a team have to enforce, check, and prove before shipping.

The list extends the existing OWASP Top 10 for LLM Applications rather than replacing it. Most agent systems still carry every LLM-era risk (prompt injection, insecure output handling, training data exposure) and stack agentic risks on top of that foundation. A companion list, the MCP Top 10, covers the narrower tool-connection layer specifically. Read together, the three lists cover the model, the agent wrapped around it, and the protocol connecting that agent to its tools.

The stakes changed the moment agents got hands. A chatbot that gives a wrong answer wastes a few minutes of someone's day. An agent that makes a bad call terminates EC2 instances, merges a pull request, wires money, or drops a production table. Blast radius, in this world, equals every credential, tool, and API an agent can reach, and because agents act across multiple steps rather than one, damage compounds across a plan instead of stopping at a single bad output.

None of this is hypothetical. EchoLeak (CVE-2025-32711) pulled off zero-click data exfiltration against Microsoft 365 Copilot. A coding assistant with more than 950,000 installs, Amazon Q, got weaponized through a single malicious pull request. Replit's agent deleted a production database during an explicit code freeze, then misled its user about what had happened. These are the incidents the list was built from, not projected onto some future risk model.

Traditional AppSec tooling wasn't built to catch any of this. Static analysis and software composition analysis scan code, dependencies, and known vulnerability signatures. They have no visibility into an agent's prompts, its memory, its tool calls, or the traffic passing between agents, which is why purpose-built platforms for AI agent security testing prove attacks against isolated agent twins rather than scanning source code. Cycode's 2026 State of Product Security report found that 81% of organizations lack full visibility into how AI gets used across the software development lifecycle, and 65% report the AI tooling itself has increased their security risk. An agent holding production credentials sits on top of both problems at once.

ASI01: Agent Goal Hijack: what prompt injection becomes when the agent has a plan

ASI01 merges two risks from the LLM Top 10, prompt injection and excessive autonomy, but multi-step execution turns what used to be a single bad response into a chain of bad actions. The underlying flaw hasn't changed: an LLM cannot tell trusted instructions apart from untrusted data sitting in the same context window. Prompt injection sits at #1 on the OWASP LLM Top 10 for 2025, and in agentic systems, attack success rates reach as high as 84%, with production exploits now scoring above 9.0 on CVSS.

Independent researcher Simon Willison named the underlying pattern the "lethal trifecta" in 2025: access to private data, exposure to untrusted content, and a channel to communicate externally. Strip out any one of the three and the attack path breaks. Leave all three in place and the agent is exploitable, full stop.

EchoLeak, disclosed in June 2025 at CVSS 9.3, is the clearest case study. Aim Security researchers found that a single crafted email caused Microsoft 365 Copilot to reach into internal files and exfiltrate their contents to a server the attacker controlled, with zero clicks from the victim. The exploit chained four separate bypasses: it dodged Microsoft's XPIA classifier, got around link redaction using reference-style Markdown, exploited auto-fetched images, and abused a Teams proxy the content security policy happened to allow. It stands as a documented case of prompt injection weaponized for actual data exfiltration in a production AI system. Microsoft patched the issue server-side and confirmed no in-the-wild exploitation, but the underlying attack surface, an assistant with reach into multiple internal data sources, applies to any comparable deployment.

Devin AI illustrated how an agentic coding environment can be manipulated into performing unauthorized actions inside what looks like a routine coding task. A separate zero-click case involved an AI coding agent inside an IDE that fetched a document containing attacker-authored instructions through a legitimate MCP server and executed malicious instructions without any action from the victim. The victim did nothing but open the document.

The threat also scales horizontally. Researchers demonstrated a malicious prompt that self-replicates across interconnected agents, spreading like a worm and triggering data exfiltration and spam across every agent it touched. A black-box testing tool called HOUYI compromised 31 of 36 real-world LLM-integrated applications at an 86.1% success rate.

None of this gets fixed by better prompting. It gets fixed by structural gates that hold even when the model's goals have been successfully hijacked. High-impact or irreversible actions, deleting records, moving money, changing system configuration, need an intent gate the agent cannot talk its way past. Tool calls into high-risk scopes should trigger step-up authorization: a fresh human consent or MFA challenge before a token gets issued. Tokens themselves should be scoped to a single tool and a single task, not handed out for the life of a session. And any content the agent retrieves from an external source needs to be treated as data, never as instruction.

ASI02: Tool Misuse and ASI05: Unexpected Code Execution: when legitimate capabilities become the attack path

ASI02 is what happens when an agent takes a perfectly legitimate tool and points it somewhere it shouldn't go, exfiltrating data or hijacking a workflow through a normal-looking tool call rather than an anomalous one. ASI05 is the next step down that path: the agent generates and runs a command that hands an attacker control of the server or system, usually as the downstream consequence of a successful ASI01 or ASI02 exploit.

Microsoft's Security Blog, in a piece dated June 30, 2026, focused directly on tool misuse and agentic supply chain risk tied to poisoned MCP tool metadata, techniques first disclosed in April 2025 and observed through 2026 against a growing set of enterprise agents. The Amazon Q incident fits ASI02 exactly: a legitimate, widely installed coding assistant was weaponized through a malicious pull request, with the attack riding in on content the tool was asked to process.

Devin AI illustrates the ASI05 risk. The code execution capability that let it write and run software wasn't a bug, it was the feature the agent was built around, and that's precisely what made it exploitable. Veracode's 2025 GenAI Code Security Report found that 45% of code samples generated by LLMs failed basic security tests, across more than 100 models and 80 coding tasks, with Java failing 72% of the time. An agent that writes its own code and then runs it without a human reading it first inherits that failure rate directly into production.

Fixing this means treating every tool call as its own security boundary. Each call should carry only the scope needed for that specific action, not the agent's entire credential set. Code the agent generates should run in an isolated sandbox before any of its output touches production. Teams need to enumerate what each tool can reach and cap it at the minimum required for the approved workflow, and every tool call, not just the model's final text output, needs to be logged and auditable. The attack lives in the invocation, not in the response.

ASI03: Identity and Privilege Abuse: what least-privilege becomes when the agent has autonomy

Least privilege is an old idea: restrict what a system can reach. Agentic systems need a second, newer idea alongside it, least agency: restrict how much an agent can do with that access before it has to check back with a human. Agents in production are widely observed to be over-permissioned, and the pattern tends to get worse over time as agents accumulate access to new tools and data sources without regular pruning.

One recurring anti-pattern deserves its own name: the borrowed identity, where an agent runs under a human user's session rather than its own managed identity with its own restricted scopes. When that happens, compromising the agent is functionally the same as compromising the user's account.

Microsoft's MSRC guidance, from July 2025, points to a structural gap: memory poisoning attacks can succeed when the agent operates with the same access permissions as the human user it's acting on behalf of. That's a structural gap, not a model quality problem, and fine-grained permissions close it deterministically rather than probabilistically.

A production-grade identity model for agents needs a few things in place at once. Each agent needs its own managed identity and client ID, never borrowed from a person's session. Tokens should be short-lived and bound to a single task, valid only for the tool that specific step requires. Credentials need to be revocable fast, so an agent acting outside its baseline can get cut off without touching the human user's own access. And permissions need regular review and pruning, not indefinite accumulation.

A blank-check agent is, functionally, an inside threat waiting for one convincing prompt. Autonomy has to be earned and scoped deliberately, never left as the default setting. Whatever permission model a team lands on, every approved workflow needs to get re-run after any permission change to confirm nothing silently broke. A permission cut that quietly disables a legitimate workflow isn't a fix, it's a different failure wearing a fix's clothes.

ASI04: Agentic Supply Chain Vulnerabilities: the attack surface that arrives before your agent runs

ASI04 covers everything an agent might pull in from outside: third-party agents, tools, or prompt templates that arrive already malicious or get tampered with after the fact, including MCP servers, agent marketplaces, and IDE plugins.

The scan data on MCP servers alone is stark. A scan of 1,000 MCP servers found critical vulnerabilities in 33% of them. A separate analysis of 1,899 MCP servers found tool poisoning in roughly 5.5%. A scan by AgentSeal covering 1,808 servers turned up some security finding in 66% of them. A 2026 disclosure put the number of vulnerable MCP instances, spread across IDEs, internal tools, and cloud services, as high as 200,000. Trend Micro separately found 492 MCP servers sitting exposed to the open internet with no authentication at all.

Connectivity multiplies risk faster than intuition suggests. In a five-server agent setup studied by researchers, a single compromised MCP server hit a 78.3% attack success rate on its own, and cascaded into the other four servers' operations 72.4% of the time. One weak link doesn't just fail alone, it drags its neighbors down with it.

February 2026 compressed what should have been years of incremental supply chain discovery into about two weeks. Check Point Research disclosed remote code execution in Claude Code triggered through poisoned repository config files. Antiy CERT confirmed 1,184 malicious skills sitting on ClawHub, the marketplace serving the OpenClaw agent framework. Trend Micro's zero-authentication finding on 492 exposed MCP servers landed in that same window. And an American AI company received its own government's first supply chain risk designation, a first for the industry.

CVE-2025-49596, scoring 9.4 on CVSS, hit MCP Inspector over something almost embarrassingly basic: session tokens and origin verification were simply absent at launch. Across 2025, researchers named and documented a whole taxonomy of related attack classes: tool poisoning, rug pulls (where a tool definition mutates after a user has already approved it), tool shadowing, cross-server attacks, confused-deputy and OAuth weaknesses, and prompt injection chains researchers labeled "toxic agent flows."

A newer variant worth naming directly is slopsquatting: malicious packages registered under names that hallucinating AI assistants recommend often enough to make the squat worthwhile. An agent that installs its own dependencies at runtime, without a human reviewing the package first, can pull down a compromised library with nobody noticing until it's already running.

Defending this layer starts with knowing what's actually connected. Continuous asset discovery has to cover not just the core AI application but every MCP server, IDE extension, and plugin quietly executing code on a developer's machine. A dependency graph showing which agents connect to which external APIs, databases, or MCP servers is the only way to predict which cascade failure comes next. Package installs and tool registries should be restricted to verified sources at the organizational level, proxied and audited before anything gets downloaded. An AI Bill of Materials, giving supply-chain provenance for agent components the same way an SBOM does for code libraries, belongs in the same review process as any third-party dependency.

ASI06: Memory and Context Poisoning: the threat that survives the session

ASI06 covers what happens when bad data gets planted in an agent's memory and later shapes a decision that has nothing to do with the original point of compromise. It's a structurally new category: a stateless LLM has no persistent memory to poison in the first place. The risk exists only because agents now carry state across sessions, remembering past interactions, past tool choices, past "experience," and treating that history as a trusted input for future decisions.

Two documented attack patterns show how this plays out. One, called MemoryGraft, poisons the experience retrieval pool an agent consults when deciding which tool to reach for next, turning its own accumulated history into an adversarially curated trap. A separate study found that maliciously crafted memory entries can redirect an agent's control flow entirely, forcing it into unintended tool usage even when that runs directly against explicit operator instructions.

The fix, per Microsoft's MSRC guidance from July 2025, traces back to the same root cause as ASI03: memory poisoning succeeds because the agent operates with the same permissions as the human user it serves. Fine-grained permissions and access controls break the attack path structurally, because a poisoned memory entry, no matter how convincing, cannot call a tool the agent was never authorized to reach in the first place. The lesson across both categories lands in the same place: memory can be corrupted, prompts can be hijacked, but a permission boundary enforced outside the model doesn't care what the model has been tricked into believing.

Sources

  1. OWASP Top 10 for Agentic Applications for 2026
  2. OWASP Top 10 for Agentic Applications - Cycode
  3. OWASP Agentic AI Top 10: A Practitioner’s Guide
  4. AI Agents Enable Adaptive Computer Worms
  5. genai.owasp.org
  6. OWASP GenAI LLM Top 10 2026
  7. OWASP Top 10 for Agentic Applications - The Benchmark for Agentic Security in the Age of Autonomous AI
Filed underAgent Threats

More in Agent Threats