Proven Exploit Requirements for AI Agent Security Findings
Unproven findings mask real agent security risks that legacy frameworks cannot see.

A finding in AI agent security is only as good as the exploit behind it, and right now, most findings don't have one. Nearly 80% of organizations have already deployed AI agents into production, and Gartner expects 40% of enterprise applications to carry task-specific agents by the end of 2026, up from under 5% in 2025. Adoption has outrun the discipline needed to tell a real threat from a plausible-sounding one, and the gap is starting to show in measurable ways.
How an agent's autonomy, tool access, and persistent state create a qualitatively different attack surface than traditional software
A conventional application has a system boundary you can draw on a whiteboard and expect it to hold. An agentic system doesn't offer that courtesy. Its attack surface grows with every tool it negotiates access to, every memory layer it populates, every message it exchanges with another agent. The boundary you drew this morning may not describe the system by the afternoon.
Most agents run as a single principal with blanket access to whatever tools and data they've been wired into. There's no compartmentalization by default, no least-privilege scaffolding unless someone built it in deliberately, ahead of time. Once an attacker steers the agent's reasoning, even briefly, they inherit everything that agent can touch.
That's why the "digital insider" framing has stuck with security researchers. An agent holds legitimate credentials. It generates activity that looks, on a log line, indistinguishable from authorized behavior, and it accumulates permissions over time, often quietly, as its task scope expands and nobody circles back to prune what it no longer needs. Security teams frequently have no visibility into that drift at all.
Then there's the part traditional models weren't built to see. System prompts and long-term memory can be poisoned once and exploited across many future sessions. A retrieval-augmented generation knowledge base isn't just something the agent reads from, it's something an attacker can write to. In multi-agent pipelines, a single compromised message can propagate across the whole chain, each agent trusting the last without question.
A systematization-of-knowledge paper by Dehghantanha and Homayoun (arxiv:2603.22928) states the underlying claim directly: agentic systems present a qualitatively different attack surface, and the risk comes from system composition, tool orchestration, and lifecycle interactions, not from prompt-level tricks in isolation. That's the structural reason a finding that would sit quietly in a backlog for a static application becomes an operational hazard the moment it's found in an agent. The blast radius scales with exactly the three things that define an agent: access, autonomy, and persistence.
Why the frameworks security teams reach for first, STRIDE, PASTA, and their peers, systematically miss how agentic attacks actually work
STRIDE, PASTA, LINDDUN, and OCTAVE were all built for systems with fixed trust boundaries and predictable component behavior. They decompose a system into parts, analyze each part against a set of categories, and produce findings that map cleanly onto those categories. That's a reasonable method for software that doesn't change its own behavior based on what it reads. Agentic systems do exactly that, and applying these frameworks to them anyway is the single most common mistake in agent security review today.
Agentic attacks break that method in three specific ways, and none of them have a home in STRIDE. Semantic state accumulation is attacker-controlled text that reshapes an agent's future reasoning without ever triggering a discrete, loggable event. Latent attacker intent persistence is an injected goal that survives across sessions by hiding in memory or in a RAG index. Cross-zone causality is an action taken in one layer of the system producing harm through entirely legitimate operations in a different layer.
Goal hijacking, listed as ASI01 in the OWASP Agentic Top 10, is the clean example. No spoofing involved, no data tampered with at rest, no privilege escalation in the classical sense that STRIDE was designed to flag. The agent simply uses the capabilities it was already granted, against the interests of the person who granted them. Run a STRIDE analysis on that, and it comes back clean, which should worry you more than a flagged finding would.
EchoLeak (CVE-2025-32711, CVSS 9.3) is the case that made this concrete. A STRIDE review of each component in Microsoft 365 Copilot, taken one at a time, would have found nothing wrong with any of them. The attack worked by moving the system through a sequence of individually legitimate states until it exfiltrated files to an external server, with no user interaction required at any point. Component-level analysis doesn't see a path like that, because the path emerges from how components interact. It's a property of the sequence, not of any single piece.
That has a direct consequence for how much weight a "finding" should carry. If the analytical method a team used cannot even model the attack path in question, then whatever that method produced amounts to an unproven risk, a category assignment and nothing more. Two teams can run the same legacy framework against the same agent and land on incompatible risk ratings for the exact same underlying exposure, and neither would be wrong by the rules of the framework they used, which is itself the indictment. Netskope's AI Risk and Readiness Report for 2026 puts a number on the resulting gap: AI tools show up in 73% of organizations, but real-time governance enforcement is in place at only 7% of them. Most organizations are making security calls with methods that can't see the attacks their own agents are actually exposed to, and mistaking the resulting silence for reassurance.
What the research community has produced to close that gap, and what it still cannot resolve
The research community hasn't ignored this. A handful of frameworks now exist specifically to model agentic risk, and each closes part of the gap that STRIDE leaves open. None of them close all of it, and pretending otherwise is the second most common mistake in this field.
MAESTRO, published by the Cloud Security Alliance in February 2025, decomposes any agentic system into seven layers running from the foundation model up through ecosystem integration. Its central claim is that the most dangerous attack paths start at Layer 1 and cascade upward through Layer 4 and beyond, a useful corrective against analysis that stops at the model layer and calls it done. CSA has already published working MAESTRO models for the OpenAI Responses API and for Google's A2A protocol.
MITRE ATLAS, at version 5.4.0, now documents 84 adversary techniques specific to AI systems. A recent expansion, done with Zenity Labs, added fourteen new agent-specific techniques and sub-techniques, including AI Agent Context Poisoning (AML.T0080), RAG Credential Harvesting (AML.T0082), and Exfiltration via AI Agent Tool Invocation (AML.T0086). Roughly 70% of the mitigations ATLAS lists map back to security controls that already exist, which is reassuring in one sense: the tooling to defend against a lot of this isn't new, it just needs to be pointed correctly.
The OWASP Agentic Security Initiative, launched in February 2025, built a taxonomy spanning agent design, memory, planning and autonomy, tool use, and deployment. Its Top 10 for Agentic Applications was released in late 2025, with ongoing development continuing into the following year. Separately, MATRA (Modeling the Attack Surface of Agentic AI Systems) takes an asset-based approach: it feeds identified assets into a business impact assessment, then builds attack trees that trace impact scenarios down to concrete attacker objectives and architecture-specific vectors. It's been demonstrated against the OpenClaw single-agent, multi-tool personal assistant architecture. NIST's Generative AI Profile (NIST AI 600-1, July 2024) adds twelve risk categories and over two hundred suggested actions on top of that.
Benchmarks have grown up alongside the frameworks. AgentDojo runs full-runtime evaluations of both benign task performance and adversarial robustness for tool-using agents. InjecAgent stress-tests indirect prompt injection specifically. Agent Security Bench and AgentHarm, formalize attacks and defenses and measure the harmfulness of agent outputs, respectively. ST-WebAgentBench, slated for ICLR 2026, extends this into safety and trustworthiness evaluation for web agents.
What this body of work gives the field is shared vocabulary, structured taxonomies, reproducible test conditions, and grounding in documented adversary behavior. What it doesn't give the field is proof. A framework sorts a risk into a category. A benchmark scores a model or an agent against a synthetic scenario built to represent a class of attack. Neither confirms that a specific attack path actually executes against a specific deployed agent, with its actual tool grants, its actual credentials, its actual configuration quirks. MATRA's own authors say as much: enterprises still lack a systematic way to trace a path from a business impact scenario to a concrete, working attack vector across the specific tools, memory, and retrieval components of a given deployed workflow. The frameworks describe what classes of problems exist. They don't tell a team whether the problem is live in their system today, and that distinction is the whole argument.
What the attack record shows about the gap between theorized risk and confirmed exploitation
The incidents that have actually happened are a better guide than the taxonomies, because each one is a specific, executed chain rather than a category on a chart.
EchoLeak, discovered by Aim Security and disclosed in June 2025, worked through a single crafted email. Hidden instructions, buried in HTML comments or in white-on-white text, caused Copilot to reach into OneDrive, SharePoint, and Teams files and exfiltrate their contents to a server the attacker controlled, without the user clicking, opening, or approving anything. The vector was Copilot's own RAG summarization pipeline, turned against itself. Microsoft patched it, but it stands as a documented, weaponized, zero-click prompt injection chain executed against a widely deployed production AI system.
CamoLeak (CVE-2025-59145), found by Legit Security and patched in August 2025 before public disclosure in October, hit GitHub Copilot Chat with a CVSS score of 9.6. Prompts hidden inside invisible markdown comments in pull request descriptions caused Copilot Chat to exfiltrate private repository secrets, and it did so through GitHub's own Camo image proxy, a component the system trusted implicitly, which is exactly what made it work.
CVE-2025-6514, disclosed in July 2025, gave a malicious MCP server full remote code execution, CVSS 9.6, against clients running the mcp-remote package. It's the first documented real-world RCE against an MCP client, and it makes a point most agent security reviews still don't account for: a third-party MCP server is its own trust boundary, distinct from the model and distinct from the agent's own tool set.
The LiteLLM supply chain breach, discovered in March 2026, involved malicious code inserted into two package versions before anyone caught it. By the time it was found, the compromised versions had spread widely across cloud environments before discovery. Sensitive contractor data was exposed across the affected environments.
Not every incident on this list involved an external attacker, and the next one shouldn't get filed under "prompt injection" just because that's the convenient bucket. Not every failure mode involves an external adversary. Documented incidents include cases where an agent's own unsupervised autonomous action, rather than any attacker, produced the harm, arguably a scarier data point than the injection cases, since there is no adversary to catch and no patch that would have stopped it.
Separately, Check Point Research's AI Security Report for 2026, drawing on a case study documented by Gambit Security, described an operator using AI-executed commands across multiple sessions, breaching nine Mexican government agencies in the process.
Look across these incidents and a small set of root causes keeps recurring in slightly different clothing: excessive agency, prompt injection, supply chain trust, and data exposure. What made each one real, rather than theoretical, is that the attack path was specific and reproducible, and it exploited the intersection of model behavior with actual tool access and actual trust relationships. Simon Willison's "lethal trifecta," described in June 2025, names the precondition precisely: private data access, untrusted content ingestion, and external communication capability, present together, are what make worst-case prompt injection possible. A finding about any one of those three legs, without a demonstration of how they combine in a given system, leaves the system's exploitability unproven. It names an ingredient. It doesn't cook the dish.
Why prompt injection is the clearest case for a proven-exploit requirement, and why theoretical findings about it are insufficient
Prompt injection sits at the top of the OWASP Top 10 for LLM Applications 2025, and no one has a complete fix for it. Frontier models from OpenAI, Google, and Anthropic remain vulnerable even after their best available defenses are applied, and that's not a knock on any one of them so much as a statement about the nature of the problem itself.
Research on MCP tool poisoning, cited across the agentic threat modeling literature, found that it evades output-based safety checks in 95% of cases. That number matters because it means a vendor's claim that a defense is "in place" tells a security team almost nothing about whether that defense actually holds against a real attack input. The gap between "we have a mitigation" and "the mitigation works" is exactly where unproven findings tend to hide, and vendors rarely volunteer which side of that gap they're on.
Part of why prompt injection resists a clean fix is that it isn't a misconfiguration you can patch away, it's an architectural property of how these systems take instructions. As the MATRA paper (Van hamme et al.) puts it, prompt injection functions the same way code injection does against a computer: it delivers new, executable instructions through a channel the system was built to accept input from. You can't patch the channel closed without breaking the feature that made the system useful in the first place.
That's precisely why a theoretical finding about prompt injection tells a security team so little on its own. Exploitability depends entirely on what tools the agent can reach, since a successful injection that can't invoke a sensitive tool isn't the same threat as one that can. Severity depends on what data and credentials the agent happens to be holding at the moment the injection lands, and persistence depends on whether the injected instruction can survive into memory, into a RAG index, or into a downstream agent that trusts the first one without checking. A finding that reads "this agent is vulnerable to prompt injection," without demonstrating the injection itself, the tool chain it activates, and the data it actually reaches, is a label. It isn't a risk assessment, and treating it like one is how backlogs fill up with noise nobody has time to sort from signal.
CrowdStrike's 2026 Global Threat Report documented threat actors injecting malicious prompts into legitimate generative AI tools at more than 90 organizations over the course of 2025. Those were operational attacks against specific, deployed configurations, not theoretical exposures sitting in a research paper. OpenAI's own response reinforces the point: the company's Lockdown Mode for ChatGPT, launched February 13, 2026, came with a public acknowledgment that prompt injection in AI browsers may never be fully patched. Once the industry accepts residual exploitability as a permanent condition rather than a bug to be squashed, the question that matters shifts. It's no longer whether this category of attack is possible in principle, but whether this specific attack succeeds against this specific deployment, and that question only has an answer once someone's actually run it.
What a proven-exploit requirement actually means in practice for agentic systems
A proven exploit, for an agent finding, has four parts, and all four have to be specific to the deployed system rather than to a general category of risk. Skip any one of them and what you have is a hypothesis wearing a finding's clothes, and too many vendor reports get away with exactly that substitution.
There's a reproducible attack input: the exact prompt, payload, poisoned document, or crafted message that triggers the unwanted behavior, written down precisely enough that someone else could run it again and get the same result. There's a demonstrated execution path: the actual sequence of tool calls, memory writes, or agent decisions the attack input produces, not a hypothesis about what it might produce. There's a confirmed impact, meaning what data got reached, what action got taken, what credential got touched, described against the agent's actual tool permissions at the moment of the attack, not its theoretical maximum. And there's isolation from production: the whole thing demonstrated in an environment that mirrors the agent's real configuration closely enough to be meaningful, without ever touching live systems or live data.
MATRA's methodology is a useful model for how this is supposed to work end to end. It starts from assets pulled out of actual system documentation, runs those through a business impact assessment, and only then builds attack trees connecting each impact scenario to concrete attacker objectives, specific techniques, and architecture-specific vectors. The tree doesn't terminate at "this class of technique exists somewhere in the literature." It terminates at a leaf node that names a feasible technique against a real, named asset in the system under review.
The isolation piece isn't a nice-to-have, it's what separates a credible disclosure from alarming noise. EchoLeak, CamoLeak, and the mcp-remote RCE were all demonstrated in controlled research environments before anyone went public, and that's precisely what made the disclosures credible rather than speculative. A finding that skips that step, however well-argued the theory behind it, is still a hypothesis. In a system where the blast radius scales with everything the agent can touch, a hypothesis is not something a security team can afford to act on as if it were proof.
Sources
- MATRA: Modeling the Attack Surface of Agentic AI Systems � OpenClaw Case Study
- AI/Agentic Threat Modeling: Securing Systems That Include Agents
- SoK: The Attack Surface of Agentic AI - Tools and Autonomy
- Agentic AI Cybersecurity Risks: How to Secure AI Agents
- cloudsecurityalliance.org
- arxiv.org
- Security Considerations for Multi-agent Systems
- Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation


