AI Agent Attack Surface Differences From Traditional API Threats
Agents expose attackers to trust boundaries that form at runtime, not design time.

Agents don't expand the traditional attack surface so much as dissolve the ground it stood on: the old assumption that inputs, trust boundaries, and execution scope can all get mapped before anything runs. Traditional API security works because you can count the pieces. Endpoints, ports, services, data flows: draw the picture once, test against it, done. Agents wreck that picture because they reason, plan, and act across systems they only discover after they've already started running. A production agent can clear static analysis, dependency scanning, and API validation without a single flag, then leak sensitive data through one prompt within minutes of going live. That's not a gap in the tooling. The tooling answers a different question than the one agents actually raise, and no amount of patching closes that gap.
How agents expand their own attack surface at runtime
A traditional API has a fixed shape. Whatever endpoints get deployed are the endpoints an attacker can go after, nothing more. Agents don't hold still like that. Every tool call, every memory write, every message passed to another agent opens new reachable state mid-task, long after deployment, not before it.
Attackers already think in kill chains, and agents play right into that habit. You don't need to break the model itself. Touch any single piece of its surroundings (a tool, a memory store, a message from another agent) and you can redirect its behavior just as well. The Cloud Security Alliance's MAESTRO framework, released in February 2025, splits this environment into seven layers: Foundation Models, Data Operations, Agent Frameworks, Deployment and Infrastructure, Evaluation and Observability, Security and Compliance, and Agent Tools and Integrations. CSA singles out Layer 7, Agent Tools and Integrations, as the place where standard security instinct fails hardest. The danger there doesn't come from any single piece. It comes from how the pieces combine across trust boundaries nobody drew on a diagram beforehand.
A follow-on to MAESTRO published in February 2026 adds a cross-layer point worth sitting with: the most dangerous attack paths start at Layer 1 and run straight through Layer 4 and beyond it. Look at any single layer alone and you miss the path entirely. Some continuous security testing tools for AI agents run attacks across the full chain in isolated twin environments for exactly this reason. That's exactly what happened with EchoLeak (CVE-2025-32711, CVSS 9.3), a zero-click prompt injection against Microsoft 365 Copilot that pulled internal files out without the user touching anything, as detailed in augmentcode's agentic threat modeling guide. Checked piece by piece, the system looked clean. The hole only shows up once someone walks the whole path start to finish, and an inventory-based model misses it by design: it counts parts instead of following the route an attacker actually takes.
The semantic layer: where inputs become instructions without any code change
Classic injection attacks, SQL injection, command injection, work by fooling a parser through syntax. So the defenses sit at that same syntactic layer: clean the input, escape the characters, check the schema. Prompt injection skips past all of that. A large language model has no reliable way to tell a trusted instruction apart from untrusted data sitting in its own context window. The attack lives in meaning, not in syntax, and no network firewall or application-layer filter can reach it, no matter how well either is tuned.
OWASP ranks this LLM01 on its Top 10 for LLM Applications 2025, and attack success rates run anywhere from 50 to 84%, depending on setup and how many tries an attacker gets. Indirect prompt injection makes things worse, because the bad instruction doesn't have to come from the attacker's own interface at all. It can sit inside an email, a web page, or a tool's output, waiting for a victim to trigger it during some ordinary task. The SafeClawArena benchmark, run in June 2026, recorded a 70% attack success rate against production-grade agents. Researchers and threat intelligence teams have separately flagged wide campaigns embedding malicious injections in ordinary public web content, targeting AI systems across public sites.
On February 13, 2026, OpenAI rolled out Lockdown Mode for ChatGPT Enterprise and education, healthcare, and teacher accounts, and said outright that prompt injection in AI browsers "may never be fully patched," per vectra.ai's analysis. No frontier model, not OpenAI's, not Google's, not Anthropic's, comes out clean once every known defense gets applied. Pick a side on this one: defense in depth isn't a best practice here, it's the only option left standing once you accept that perimeter defenses and input checking aim at the wrong layer entirely. The problem lives in how the model reasons, not in the shape of what gets fed to it.
Trust boundaries in multi-agent systems are runtime properties, not design-time facts
API security assigns trust once, at design time. Service A gets permission to call service B, that relationship gets written into config, and a perimeter control enforces it forever after. Multi-agent systems don't get that luxury, because trust moves through message passing: a sub-agent's output becomes the orchestrator's trusted input, full stop, no further check behind it.
Johann Rehberger documented cross-agent privilege escalation in September 2025: a compromised sub-agent hands the orchestrator a manipulated output, the orchestrator treats it as trustworthy, and the attacker walks straight out of a low-privilege sandbox into a high-privilege execution context. Researchers now call the resulting spread "prompt infection," malicious instructions passed agent to agent, building persistence across tasks and triggering unauthorized actions well downstream of the original break-in. The OWASP Agentic Top 10, published December 9, 2025 with input from more than 100 industry practitioners, names this pattern ASI01, goal hijacking, and it doesn't fit into any STRIDE category cleanly. No spoofing happened. No data got tampered with at rest. No privilege got escalated in the classical sense. The agent just used its own legitimate powers against the person who deployed it, and that's worse than any classical category, because there's no broken rule to point to afterward.
A June 2026 paper by Krishna Mohan and Guda Nagavenkata Srinivasa lays out six core agentic threat categories: prompt injection, identity and authorization, action auditability, tool abuse, data residency, and boundary policy enforcement. It finds that frameworks like LangGraph and Google's ADK leave every one of these controls up to whoever deploys the agent. There's no default, and that's the part worth sitting with: security here isn't a switch the framework flips on, it's homework left for the deployer, and most deployers don't even know the assignment exists. The MATRA framework, accepted for DeMeSSAI 2026 alongside IEEE EuroS&P 2026 in Lisbon, models how controls like network sandboxing and least-privilege access shrink blast radius without killing off the underlying attack class. Trust in an API system gets baked into how the thing is deployed. In an agentic system, trust is a property of each message, checked or not, in flight right now.
Memory and state persistence create attack surfaces that survive individual sessions
Traditional APIs are mostly stateless. A compromised request stays contained to that one request, with no built-in path for the damage to carry into the next call. Agentic systems break this on purpose: short-term working memory, long-term memory stores, RAG knowledge bases, all of it persists context across sessions because that's what makes an agent useful over time instead of starting from zero each run. That same design choice turns memory into a liability rather than just a feature.
Memory poisoning corrupts an agent's stored context with false or misleading information, so every future session inherits reasoning that's already compromised, as security researchers have documented. RAG index poisoning is a variant on the theme: plant a handful of bad documents in the retrieval corpus, and every future retrieval gets steered off course from there on. One write, indefinite read-time damage. Temporal persistence is its own threat domain, something STRIDE has no category for at all. A separate MATRA paper (arxiv 2605.10763) notes that system prompts and long-term memory stores make good targets for long-horizon influence campaigns, ones that compromise an agent slowly across many sessions instead of in one shot.
Get this distinction wrong and the fix lands nowhere. An attack on the agent, corrupting what it remembers, is a different problem from an attack through the agent, using that poisoned memory to reach some downstream database, and each needs its own fix. Confusing the two is the mistake most teams make first, and it's an expensive one: patching the wrong layer leaves the actual hole wide open. An agent that looks perfectly secure today may already be carrying attacker-controlled context into tomorrow's session, a threat class with no clean classical equivalent to reach for.
The magnitude difference: why a compromised agent is not a compromised user account
A compromised user account stays bounded. It reaches exactly what that account could reach, nothing further. Agents don't respect that fence, mostly because nobody built one for them to respect in the first place. Obsidian Security's analysis of prompt injection found agents move 16 times more data than human users do. Treating a compromised agent like a stolen password, the reflex most security teams default to, gets the scale of the thing wrong by an order of magnitude, and that's the mistake worth naming directly: a stolen password buys an attacker one account, while a compromised agent buys them every system that agent was ever trusted to touch, which is usually the whole environment, not a corner of it.
IBM's Cost of a Data Breach Report 2025 found 13% of organizations had already suffered a breach involving an AI model or application, and 97% of those lacked proper AI access controls. Netskope's AI Risk and Readiness Report 2026 found AI tools deployed at 73% of organizations, but real-time governance enforcement running at only 7% of them. That gap, deployment racing ahead of control by a factor of ten, is the whole story. Agentic systems read and write data, call outside APIs, and run multi-step workflows with limited human oversight, which opens the door to large-scale data theft, supply-chain compromise, and behavior nobody planned for going in.
None of this is hypothetical. GitHub Copilot had a confirmed remote code execution flaw (CVE-2025-53773) with a working proof-of-concept, patched in August 2025. Claude Code had a DNS exfiltration bug (CVE-2025-55284). Researchers showed that a webpage could inject instructions into Anthropic's Claude Computer Use and redirect its behavior in unauthorized ways. Slack's AI agent leaked private channel information after a prompt injection delivered through a public channel. Perplexity's Comet agent got redirected by website injections into leaking user data to servers the attacker controlled. Separately, one SoK paper (arxiv 2603.22928) found 19 remote code execution flaws with working exploits across 11 different agent frameworks, all stemming from weak tool schemas and validation: the classical software layer sitting right under the semantic one, both exposed at once. The agent's tool access, its memory, its role coordinating other systems: every one of those acts as a multiplier on whatever gets through the front door.
Why the frameworks built for traditional threat modeling cannot be retrofitted
STRIDE, PASTA, and DREAD all lean on three assumptions: stable system boundaries, deterministic behavior, and components that can be checked one at a time. Agentic systems break all three, consistently, by nature, not as some edge case that slipped through the cracks.
Three failure modes show up specifically here, and none of them have a home in the old frameworks. First, semantic state accumulation: no STRIDE category asks what happens when an agent's future reasoning depends on attacker-controlled text it's been carrying since three turns ago. Second, cross-zone causality: an EchoLeak-style attack runs input, then retrieval bias, then a planning goal shift, then tool invocation, then aggregated exfiltration, and STRIDE treats each arrow in that chain as its own separate check rather than one continuous path worth tracing end to end. Third, abuse of legitimate functionality: no spoofing, no broken authentication, no tampered binary anywhere in the chain, every step works exactly as designed, and STRIDE can flag misuse of a single component but has no way to model misuse stacked across several of them at once.
Purpose-built frameworks have started closing these gaps, and they're worth taking seriously precisely because they were built for this problem rather than adapted from an older one. MITRE ATLAS, now at version 5.4.0, catalogs 16 tactics, 84 techniques, 56 sub-techniques, 32 mitigations, and 42 case studies. A collaboration with Zenity Labs added 14 new techniques aimed squarely at agent security, including AI Agent Context Poisoning (AML.T0080), Modify AI Agent Configuration (AML.T0081), RAG Credential Harvesting (AML.T0082), and Exfiltration via AI Agent Tool Invocation (AML.T0086). About 70% of ATLAS mitigations map back to security controls that already exist, and the data ships in STIX 2.1, so it plugs straight into a pipeline instead of sitting in a PDF somewhere. OWASP's Agentic Security Initiative guide, published February 2025, lays out a threat taxonomy across Agent Design, Agent Memory, Planning and Autonomy, Tool Use, and Deployment and Operations. Its Top 10 for Agentic Applications, published December 9, 2025, is a separate document from the LLM Top 10, built to map threats onto autonomous agent architectures directly, and it catalogs ten risk categories, ASI01 through ASI10, covering agent identity and privilege abuse, agentic supply chain vulnerabilities, and unexpected code execution among them. NIST AI 600-1, the Generative AI Profile from July 2024, translates the broader NIST AI Risk Management Framework into 12 risk categories specific to LLM systems. MATRA takes an impact-first approach, using attack trees to connect CIA-based scenarios to vectors specific to a given architecture, and it's been demonstrated against OpenClaw ahead of its DeMeSSAI 2026 presentation.
A SoK paper out of the University of Guelph and Aalborg University (arxiv 2603.22928), pulling together evidence from 2023 through 2025, argues for attacker-aware metrics like Unsafe Action Rate and Privilege Escalation Distance, since a binary pass-or-fail test can't describe an attack chain that's probabilistic, stateful, and multi-step by nature. One MCP-specific finding stands out: tool poisoning slips past output-based safety checks 95% of the time, per augmentcode's guide. The old habit of catching problems after the fact, a reflex borrowed straight from API security, just doesn't hold at this layer. There's a timing problem stacked on top of all this, too. Agent-authored code can show up in seconds, while manual threat modeling takes hours, and in that gap an agent can change code, call tools, write to memory, and update dependencies before anyone's even run a repository checkpoint.
What a purpose-built security approach for agents must actually verify
Each structural break above points to something specific that needs checking, not a general principle to keep in the back of your mind. Dynamic surface expansion means verification has to cover the attack surface an agent builds while it runs, not the surface someone mapped at design time and called finished. Semantic injection means testing has to happen at the reasoning layer itself, since static analysis and syntax scanning are blind to failures that live entirely in how a model reads meaning rather than how it parses characters.
Dynamic trust boundaries mean every message passed between agents gets treated as a possible boundary crossing, not waved through like some internal call between two services that trust each other by default. Memory persistence means security posture can't get measured session by session; persistent memory and RAG stores need their own integrity checks, ones that hold across sessions rather than resetting with each new one. And blast radius multiplication means the stakes of getting any of this wrong scale with what the agent can touch, not with what a single compromised account could have reached under the old model.
None of these four requirements fit onto a checklist built for API security. Fixed scope, deterministic behavior, stable trust, stateless requests: those are the assumptions that checklist depends on, and not one of them survives contact with something that reasons, remembers, and acts across a surface it's still discovering as it goes. Anyone still reaching for STRIDE as the whole answer here is solving yesterday's problem with yesterday's tool, and the gap between the two only grows the longer that habit sticks around.
Sources
- MATRA: Modeling the Attack Surface of Agentic AI Systems � OpenClaw Case Study
- SoK: The Attack Surface of Agentic AI - Tools and Autonomy
- AI/Agentic Threat Modeling: Securing Systems That Include Agents
- AI Agent Security Attack Surface Map [2026 Checklist]
- The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey
- AI Agent Security: Threats, Controls, and Governance
- AI Agent Attacks in Q4 2025 Signal New Risks for 2026 | eSecurity Planet
- practical-devsecops.com


