Est.

Threat Intelligence Coverage Gaps for AI Agent Attack Patterns

Security teams lack tools to see AI agents' runtime behavior changes.

Correspondent · · 11 min read
Cover illustration for “Threat Intelligence Coverage Gaps for AI Agent Attack Patterns”
Threat Platforms · September 26, 2026 · 11 min read · 2,484 words

Threat intelligence, as a discipline, was built to answer one question: what can go wrong inside a system that stays roughly the same shape over time. AI agents don't stay the same shape. They negotiate new tools mid-task, write to memory that persists across sessions, and pass instructions between other agents in chains nobody mapped in advance. The frameworks designed for fixed systems have no vocabulary for that, and the gap between what security teams think they're watching and what agents actually do has become the defining blind spot in enterprise AI deployment.

Why traditional threat intelligence frameworks cannot see what AI agents do

STRIDE, component-level threat modeling, API gateway inspection: these tools all assume a system with fixed trust boundaries and behavior you can predict in advance. Spoofing, tampering, repudiation, information disclosure, denial of service, elevation of privilege. Six categories, each built around the idea that a system's parts and permissions stay put long enough to draw a diagram of them.

AI agents break that assumption at every layer. The attack surface grows every time an agent negotiates a new tool, writes to a memory store, or exchanges a message with another agent, and it does this at runtime, not at design time. An API gateway can confirm a request is well-formed, properly authenticated, and inside its rate limit. It has no way to tell whether the natural-language prompt riding inside that request is trying to override the system's instructions, pull data out through a tool call, or quietly steer the agent toward something nobody authorized.

It's a category error, not a problem more rules or more sensors will fix. It's a category error. The frameworks weren't built to model an object that changes its own shape while it runs, so asking them to cover agentic behavior is asking a ruler to weigh something.

How governance deployments have outrun security visibility

Diagram: Governance Has Raced Ahead of Security Visibility. Visualizes: Show the stark magnitude contrast between three figures from the article: 73% of organizations now run AI tools, but only 7% have real-time governance enforcement, and only 16%…

The numbers make the mismatch concrete. AI tools now run inside 73% of organizations, but real-time governance enforcement, the kind that actually watches what an agent does as it does it, reaches only 7% of them, leaving most of the enterprise world running agents with no live check on their behavior. That's a substantial lag. That's most of the enterprise world running agents with no live check on their behavior.

Only 16% of organizations say they effectively govern which core business systems their AI can touch. And the consequences show up in the breach data: 89.5% of organizations reported at least one generative AI-related security incident in the preceding twelve months, up from 75.1% in the prior year's figures. Thirteen percent reported breaches specifically involving AI models or applications. Ninety-seven percent lacked proper access controls over what those models could reach.

Combining those factors produces a pattern that has nothing to do with sophistication of attack. Organizations are deploying agents faster than they're building the visibility to watch them, and the breach rate is simply catching up to that gap.

The four structural blind spots that STRIDE-era methods share

Four specific holes appear wherever a component-level threat-modeling framework gets applied to agents, and each one traces back to the same root cause: the framework assumes fixed structure where agents have none.

Goal hijacking is the first, and it doesn't fit anywhere in STRIDE's six categories. The OWASP Agentic Top 10 names it directly as ASI01, Agent Goal Hijack. Nothing gets spoofed. No data gets tampered with at rest. No privilege gets escalated. The agent just uses the capabilities it already has, legitimately, toward a goal the attacker chose instead of the one it was given. A STRIDE assessment against a hijacked agent comes back clean, because STRIDE was never built to notice a system doing what it's authorized to do, just for the wrong reason.

Second, component analysis misses attack paths that only exist across components. EchoLeak, the zero-click prompt injection against Microsoft 365 Copilot, pulled internal files out without any user interaction, and every individual piece of that system passes a per-component STRIDE review. The vulnerability becomes visible only when someone traces the full path: injected content, through retrieval, through tool invocation, to exfiltration. Any single link in that chain looks fine on its own.

Third, trust boundaries in agent systems aren't fixed, they're probabilistic. An instruction telling an agent "never share sensitive data" is a suggestion baked into a prompt, not an enforced policy the way a firewall rule is enforced. Tool poisoning attacks get past output-based safety checks roughly 95% of the time, because the check is looking at what the model says it will do, not verifying what it actually can do. Static trust-boundary diagrams can't represent a boundary that reshapes itself every time the agent negotiates a new tool.

Fourth, persistence through memory has no classical analog. The ATFAA framework treats temporal persistence as its own threat domain precisely because STRIDE doesn't cover it. Memory poisoning through tampered logs or corrupted embeddings doesn't map to tampering, disclosure, or any of STRIDE's six buckets. It's a threat class that only exists once a system has memory that persists and updates on its own, which agents have and static systems don't.

What three purpose-built frameworks cover that legacy models cannot

Three frameworks have emerged that were built for agents from the start, and each one tracks attack paths across a system, not just risks sitting inside individual components.

MAESTRO breaks any agentic system into seven layers: Foundation Models, Data Operations, Agent Frameworks, Deployment and Infrastructure, Evaluation and Observability, Security and Compliance running vertically across all of them, and Agent Tools and Integrations. That last layer, Layer 7, is where conventional security instinct breaks down hardest, because the risk there comes from how tools compose across organizational and trust boundaries, not from any flaw in a single tool. The framework's cross-layer design holds that the most dangerous attack paths start at foundational layers and cascade forward through deployment and beyond, which is precisely the kind of multi-step path that component-level STRIDE analysis can't see by design. A February 2026 update adds Implementation and Continuous Monitoring coverage, and an industry security group has already published working layered threat models against the OpenAI Responses API and Google's A2A protocol.

MITRE ATLAS takes a different approach: it starts from documented adversary behavior rather than architectural layers. It now contains 84 documented techniques, and a collaboration with Zenity Labs added agent-specific entries including AI Agent Context Poisoning, Modify AI Agent Configuration, RAG Credential Harvesting, and Exfiltration via AI Agent Tool Invocation. Every technique ties back to a real case. The poisoned Postmark MCP server used for email exfiltration, the Bing Chat indirect prompt injection incident, LAMEHUG malware attributed to the Russian state-backed group APT28. Roughly 70% of ATLAS's recommended mitigations map onto security controls that already exist inside typical enterprise stacks, and the data ships in STIX 2.1 format so it can feed directly into detection pipelines without a translation layer. OWASP's LLM01 prompt injection category maps to three separate ATLAS techniques spread across two tactic stages, so treating prompt injection as one technique instead of a staged sequence misses how the attack actually unfolds.

The OWASP Top 10 for Agentic Applications 2026, built with input from more than 100 industry practitioners, is the first formal taxonomy built specifically around what autonomous agents do wrong. It covers goal hijacking, tool misuse, identity abuse, memory poisoning, cascading failures, and rogue agent behavior, the exact list of failure modes STRIDE has no words for.

What ties the three together is the path. Each one is built to trace how a failure introduced at one layer turns into an exfiltration event several layers downstream, which is exactly the kind of connective analysis that component-by-component review was never designed to do.

Prompt injection as the clearest example of a coverage gap in practice

Prompt injection isn't a bug that a patch will close, because the root cause is architectural. Large language models process the system prompt, the user's input, and any text pulled in from external sources as one continuous stream of tokens. There's no reliable boundary between an instruction and data inside that stream, so a malicious instruction buried inside a retrieved web page or document reads to the model exactly the same as a directive written by its own developers.

The scale is not small. Research has recorded attack success rates as high as 84% against agentic systems. InjecAgent found GPT-4 running under a reasoning-and-acting prompting pattern vulnerable 24% of the time against baseline attacks, and that number climbed to roughly 47% once researchers used enhanced attack techniques. Google researchers tracked a 32% rise in malicious prompt injection payloads embedded in web content between November 2025 and February 2026.

An agent becomes exploitable the moment it has access to private data, exposure to untrusted content, and a way to send information out to the world, all three at once; researchers identify this combination as the core prerequisite for exploitation. Removing any one leg causes the attack path to collapse. This single structural fact explains nearly every major prompt injection incident on record, and it hands practitioners something more useful than a growing list of exploits to patch: a diagnostic question to ask about any agent before it ships. Does it have all three properties? If yes, it's exploitable regardless of what safety instructions sit in its system prompt.

Goal hijacking widens the damage further once agents start working in pipelines. Redirect one agent's objective in a multi-agent chain, and that redirection can propagate forward, instructing downstream agents, poisoning shared memory stores, or manipulating whatever orchestrator is coordinating the group. The blast radius scales directly with how much autonomy and tool access the compromised agent was given.

Diagram: The Three-Leg Prerequisite for Prompt Injection Exploitation. Visualizes: Visualise the three structural conditions that must co-exist for a prompt injection attack to succeed: (1) access to private data, (2) exposure to untrusted content…

How the Model Context Protocol opened a new attack surface that existing coverage misses entirely

The Model Context Protocol, MCP, made the relationship between an AI model and external tools explicit by standardizing how a model calls out to APIs and data sources. API security is genuinely the foundation everything else in AI security rests on, but a request that's syntactically valid, properly authenticated, and inside its rate limit can still carry a natural-language payload built to manipulate the agent reading it.

The specific gap sits in MCP tool descriptions, the short blocks of text that tell an agent what a tool does and how to call it. These descriptions get read by the model but rarely get shown to the human user at runtime. An attacker can therefore bury a hidden instruction inside a tool description, and the model will follow it without anyone watching the interface ever seeing it happen. That instruction can direct the agent to take an unauthorized action, pull data out through a side channel, or suppress the notification that would otherwise flag its own behavior.

The scale of exposure is stark. By early 2026, roughly 8,000 MCP servers sat exposed on the public internet with no authentication in front of them. Security researchers documented more than 30 distinct vulnerabilities across the MCP ecosystem inside a single 60-day window. As of early 2026, at least seven confirmed high- or critical-severity CVEs span major MCP-integrated platforms. That's a new class of infrastructure standing largely unguarded at internet scale.

Why multi-step attack paths and escalation sequences fall outside current detection scope

Researchers have begun mapping escalation sequences for agentic attacks in ways that help locate where most enterprise threat intelligence programs currently sit, and where their coverage runs out.

Multi-step escalation sequences range from indirect prompt injection producing persistent state changes, through RAG-based poisoning and indirect exfiltration incidents, up to full agentic exploitation through tool chains and MCP servers, where a compromise cascades across systems, the pattern behind tool poisoning attacks and documented exploitation through agentic tool chains. The most advanced pattern, still emerging, covers multi-agent system attacks capable of cascading across organizational boundaries entirely, through exploitation of the trust relationships agents extend to each other.

That last level points at the real detection problem. Once the actors involved are agents talking to other agents rather than humans talking to systems, the attack surface is the network connecting the systems, not merely the sum of each individual system's weaknesses. Detection logic built for the former case has no natural way to reason about the latter.

EchoLeak makes the point concrete one more time. Component-by-component review finds nothing wrong anywhere in the chain. Only analysis that traces the full path, from the moment an instruction enters, through retrieval and tool invocation, to the point data leaves the system, can reveal the attack. That's a detection architecture that was never built to watch for a sensor failing to fire. It's a detection architecture that was never built to watch for a path.

And the reliability engineering that makes agents useful only deepens the exposure. Getting an agent to behave predictably takes orchestration layers, guardrails, evaluation harnesses, fallback logic, and multiple points where systems integrate with each other. Every one of those layers, necessary as it is for making the agent work reliably, is also a new surface an attacker can aim at. The more engineering effort goes into making an agent dependable, the larger the attack surface tends to grow alongside it.

Why standard testing and red-teaming methods cannot validate agent behavior before it ships

Conventional red-teaming assumes a target that behaves the same way twice. Run the same input through a web application or an API endpoint and, barring a code change, the output is deterministic enough that a passed test today still means something tomorrow. Agents don't offer that guarantee. The same prompt, run against the same agent, can produce different tool-call sequences, different memory writes, and different downstream actions depending on what context the agent happened to retrieve or what state its memory was in at that moment.

That non-determinism means a red-team exercise that finds no vulnerability on Tuesday says very little about what the agent will do on Thursday, once its memory has accumulated new context or a connected tool's description has changed. Testing a fixed set of prompts against a fixed set of expected outputs, the standard shape of most security testing programs, catches the failure modes that look like the failure modes people already know to test for. It has no natural way to catch a goal hijack, where the agent performs an authorized action toward an unauthorized end, because nothing about that action looks anomalous from the outside.

Closing this gap means testing paths instead of endpoints: tracing what happens when a tool description changes, when memory gets written to by an untrusted source, when one agent in a chain gets nudged off its stated goal. That's a fundamentally different discipline from the request-response testing that security teams have relied on for the last two decades, and building it out is the actual work standing between where agent governance is today and where the breach data says it needs to be.

Sources

  1. AI/Agentic Threat Modeling: Securing Systems That Include Agents
  2. SoK: The Attack Surface of Agentic AI - Tools and Autonomy
  3. cloudsecurityalliance.org
  4. systemprompt.io
  5. labs.cloudsecurityalliance.org
Filed underThreat Platforms

More in Threat Platforms