Offensive Security Operations Augmented by AI Agents
Models are just one-third of your AI attack surface, and attackers know it.

The attack surface security teams need to reason about is roughly three times larger than a model inventory suggests, and most of the offensive risk sits in the two-thirds that inventory misses entirely. Snyk data from enterprise environments, reported by beam.ai, puts the model itself at only about a third of the picture; the rest is agent frameworks, MCP servers, vector databases, retrieval pipelines, datasets, and third-party tools. Most of the AI packages and tools running in these environments come from outside the organization, so the bulk of the agent stack is code nobody inside the company wrote and nobody inside the company fully controls.
Lineage makes this worse. Roughly half of organizations running AI models can't trace those models back to the datasets used to train or fine-tune them, which turns supply chain traceability into a live opening for anyone looking to exploit it. Counting models to gauge AI risk works about as well as counting front doors to secure a building: it tells you almost nothing about the windows, the loading dock, or the maintenance tunnel underneath.
Five components make up the surface that model inventories skip over: agent frameworks, MCP servers, vector databases and retrieval pipelines, datasets and model lineage, and third-party packages. None of these show up on a model list, and each one carries its own permissions, its own code paths, and its own chances for something to go wrong. Adoption of agentic architectures, built on AI agents, MCP servers, or both, has already reached nearly half of organizations, and that adoption curve is running well ahead of the governance needed to manage it. The gap between what's deployed and what's actually watched is where offensive operations, both defensive red teams and real attackers, now spend their time.
Autonomous agency expands the attack surface for attackers and defenders alike
Size alone doesn't capture what makes this threat different. The move from a large language model to an agent is not a matter of degree: it hands the attacker a new kind of target, one that holds credentials, remembers past sessions, and can cause harm that outlasts the conversation that triggered it. Agentic systems, by contrast, carry read-write API and database access, retain persistent memory, hold elevated permissions by design, and require behavioral detection to catch misuse. Traditional LLMs stay read-only, hold only session-based memory, operate within a sandbox scope, and produce misbehavior that pattern-detection can catch.
Agents are granted broad permissions so they can act with minimal human oversight, and that design choice is what attackers rely on. It isn't a flaw in one company's implementation; it's the feature the entire category is built around. Security researchers have a name for the resulting exposure: the confused deputy problem. An attacker doesn't need to breach the network directly when a trusted agent can be talked into acting on the attacker's behalf, carrying its own legitimate permissions into whatever the attacker wants done.
Practitioner opinion has moved fast on this point. A Dark Reading poll cited by kiteworks.com found that nearly half of security professionals rank agentic AI as the top attack vector for 2026, ahead of deepfakes and every other category on the list. That ranking reflects a genuine shift in how the field sees the threat, not just headline anxiety. The tooling hasn't caught up: SIEM and EDR platforms were built to flag anomalies in human behavior, and an agent carrying out an attacker's instructions across thousands of individual steps can look entirely ordinary to systems tuned for a human operator's pace and pattern.
The threat taxonomy: frameworks, memory, planning, tools, and deployment
Several reference models now exist to map where agentic risk actually concentrates, giving practitioners a structure to work from instead of an ad hoc list of worries. The OWASP Agentic Security Initiative organizes the surface into Agent Design, Agent Memory, Planning and Autonomy, Tool Use, and Deployment and Operations, extending the OWASP Top 10 for LLMs to cover risks specific to autonomous systems. MITRE's ATLAS matrix offers a parallel reference practitioners can run alongside it, and the Cloud Security Alliance's MAESTRO framework circulates widely too, sometimes treated as its own model and sometimes folded into ASI's architectural threat-modeling layer.
A newer framework, MATRA (Modeling the Attack Surface of Agentic AI Systems), pushes this further by tying impact directly to architecture. The authors tested it against OpenClaw, an open-source personal assistant that runs as a single agent with access to multiple tools: messaging channels, shell access, file I/O, web browsing, and persistent memory. The case study captures both deliberate adversarial manipulation and the non-adversarial case where an LLM's own unexpected behavior combines with autonomous tool execution to produce a high-impact outcome.
MATRA and the adjacent frameworks expose an architectural blind spot: automated checks can validate a call's format and authentication, but none of them can inspect the intent riding inside it. An API gateway can confirm that a call is well-formed, properly authenticated, and within its rate limit, but it has no way to tell whether the natural-language prompt riding inside that call is trying to override system instructions, pull data out, or steer the agent somewhere it shouldn't go. That gap is what separates AI security from ordinary API security bolted onto a new kind of endpoint. It also explains why point-in-time testing falls short: system prompts, long-term memory, and RAG knowledge bases are persistent components that an attacker can target for slow, long-horizon influence, compromising a system across sessions rather than in one visible strike.
Prompt injection as an exploitation discipline, not a prompt-safety problem
Prompt injection sits at the center of nearly every technique described above, and the field now treats it as a structural exploitation discipline rather than a content-moderation problem to be tuned away. OWASP contributor Ariel Fogel, speaking at Infosecurity Europe 2026, called it an unresolved problem: large language models process every input, system prompt, user query, and retrieved content alike, as one continuous token sequence, with no reliable mechanism to enforce privilege boundaries between those sources, which is the mechanical root of the issue. Prompt injection functions the same way code injection does against a computer: it delivers new instructions the system will execute, except here the target is a model rather than a runtime. MATRA's authors are blunt about the ceiling on current defenses: they can lower the odds of a successful injection or catch one after the fact, but none of them close the door entirely.
Simon Willison's "lethal trifecta" identifies three conditions that make an agent exploitable. Three conditions have to be present together: access to private data, exposure to untrusted content, and a channel to communicate externally. An agent with all three is exploitable by definition; strip away any single one of the three and the attack path collapses.
Attackers overwhelmingly lean on social engineering to get the injection through. Unit 42 research found that most observed cases of indirect prompt injection use an authority-override frame, phrases like "this is a security update," or language that dresses the malicious instruction up as a routine system task. Check Point researchers Yarden Porat and Shahar Tal extended the picture at Black Hat USA 2026, showing that LangChain, CrewAI, AutoGen, and Semantic Kernel all carry exploitable logic in their core runtimes, memory stores, planning loops, and serialization layers, independent of what tools the agent has access to.
The sharpest illustration of what makes agentic injection different from a chatbot jailbreak is memory poisoning. Lakera AI research demonstrated the technique in production systems: an attacker plants false instructions into an agent's long-term memory through a poisoned data source, and the agent recalls and acts on those instructions days or weeks later, in sessions that have nothing visibly to do with the original attack. Asked about it afterward, the agent defends the planted belief as its own. A sleeper agent sits inside a production system long after the session that planted it ends.
Agent frameworks and the supply chain as the delivery mechanism
Injection describes how an attacker talks to an agent once it's running. Getting the payload into that agent before it ever reaches production is a separate problem, and the agent framework and its dependency tree have become the primary route in. The clearest example so far is LiteLLM. In March 2026, a package on PyPI named hackerbot-claw carried a backdoor and sat live for three hours, during which nearly 47,000 downloads happened. LiteLLM serves as the language-model gateway for CrewAI, DSPy, Microsoft GraphRAG, and dozens of other frameworks; anyone who pulled an update during that window pulled in an autonomous attack bot named hackerbot-claw. A single compromised package propagated across an entire ecosystem at once, simply because so much of that ecosystem depends on the same gateway.
Individual product vulnerabilities show the same pattern at smaller scale. CVE-2026-22708, found in Cursor, let an attacker use shell built-ins like export and alias to bypass an allowlist meant to restrict which commands could run; once the shell environment itself was poisoned, ordinary allowlisted commands such as git branch would carry out arbitrary payloads. The allowlist didn't just fail to stop the attack, it actively auto-approved the exact commands the attacker needed. CVE-2025-59532, found in OpenAI's Codex CLI, worked differently but arrived at a similar result: a bug in sandbox configuration logic let the model's own generated current working directory get treated as the sandbox's writable root, which quietly erased the workspace boundary the sandbox was supposed to enforce.
At Black Hat USA 2026, Francesco Montorsi and colleagues from Zenity, in "Promptware EOD: Skillful Agent Detonation," treat the entire AI agent supply chain as a malware delivery vector, with payloads hiding in skill markdown files, rug-pulled MCP servers, misaligned models, and weaponized posts. The team open-sourced a detonation chamber built for this kind of artifact, using agent honeypots and behavioral analysis, described as the first serious attempt at building malware analysis tooling for the AI agent ecosystem. Separately, NVIDIA researchers Bar Lanyado and Eliya Cohen presented WASP-OS, a fine-tuned open-source model that matches frontier commercial models at exploiting AI agents, at a fraction of the cost and without needing API access to any commercial provider. An attacker no longer needs a frontier model account to run frontier-grade exploitation, a democratization of offensive capability.
Unvetted open-source MCP servers, fast "vibe coding" deployment cycles, and a dependency tree made mostly of external packages combine, leaving organizations carrying attack surface they didn't build and can't fully audit.
What real incidents involving agent exploitation look like
None of this is hypothetical anymore. A mid-market manufacturer deployed an agent-based procurement system in the second quarter of 2026; by the third quarter, Stellar Cyber reported, attackers had compromised the vendor-validation agent through a supply chain attack on the AI model provider itself. The compromised agent started approving purchase orders from shell companies the attackers controlled, and the fraud went unnoticed until inventory counts dropped sharply enough to force an investigation, by which point millions of dollars in fraudulent orders had already cleared. The root cause traced back to a single compromised agent in a multi-agent system, whose false approvals cascaded downstream through every process that trusted it.
The Agents of Chaos study, run by Shapira and colleagues in 2026, gives a broader view of what goes wrong when agents get real access. Twenty researchers spent two weeks interacting with autonomous agents deployed live, with persistent memory, email, chat, filesystem, and shell access. They documented eleven distinct case studies, covering agents that complied with instructions from people who weren't their owners, disclosed sensitive data, took destructive actions on live systems, spoofed identities, and passed unsafe behavior on to other agents in the same environment. Eleven separate failure modes, from twenty researchers, in two weeks, against systems that were live rather than staged for the test.
Scale confirms the pattern holds industry-wide. Cloud Security Alliance and Token Security, in a report titled "Autonomous but Not Controlled: AI Agent Incidents Now Common in Enterprises," found that a majority of organizations have experienced at least one cybersecurity incident caused by an AI agent in the past year. A retail-focused red team engagement, presented at Black Hat USA 2026 by Netanel Rubin and Dan Avraham of Rein Security under the title "Bye Bye AI: How We Hacked the AI Shopping Assistant of a Top 3 US Retailer," bypassed the assistant's guardrails, pulled out customer data, and manipulated live purchase flows against a production system belonging to one of the country's largest retailers. What connects the manufacturing fraud, the Agents of Chaos case studies, and the retail engagement is that none of them stayed contained at the point of entry. Compromise spread through agent memory, through trust relationships between agents, and through downstream tool calls, often sitting dormant long enough that the original injection point was invisible by the time anyone went looking for it.
Why standard testing methods break on agentic systems
That dormancy and propagation pattern is what conventional security testing isn't built to catch. Static and dynamic application security testing both assume that a given input produces a consistent, repeatable outcome, an assumption that holds for ordinary software and falls apart for agents. Large language models generate output from a probability distribution rather than a fixed function, and agents chain those outputs together into loosely coupled execution graphs where a small change in context can send the whole sequence somewhere different. Feeding the same prompt into the same agent twice, in slightly different conversational contexts, can produce runs that diverge in ways that have nothing to do with a bug in the traditional sense.
Deterministic pass-fail logic, the backbone of SAST and DAST, has no reliable way to characterize a surface that behaves like this. Catching what actually goes wrong in an agentic system takes behavioral analysis and execution monitoring that follows a request across every prompt, every agent, and every tool call it touches, instead of a single test checking whether one function returns the expected value. That's the shift testing methodology now has to make, and it's the same shift the rest of this piece has been tracing: from securing a fixed artifact to monitoring an actor that keeps acting long after the code review is over.
Sources
- AI Agent Security 2026: Attack Surface Is 3x Your Models
- MATRA: Modeling the Attack Surface of Agentic AI Systems -- OpenClaw Case Study
- Top Agentic AI Security Threats in Late 2026
- AI Agents Take Center Stage at Black Hat USA 2026 | Straiker
- Agentic AI Attack Surface: Why It’s the #1 Cyber Threat of 2026 and How to Secure It
- 2026: The Year Agentic AI Becomes the Attack-Surface Poster Child
- Prompt injection still drives most agentic AI security failures in production - Help Net Security


