Guardrail Bypass Techniques Against AI Agent Runtime Controls
Attackers now have documented ways to bypass every layer of AI agent security controls.

Runtime guardrails for AI agents fail in specific, repeatable ways, and the techniques that break them are now documented well enough to name. Cisco's State of AI Security 2026 report found that only 29% of organizations feel ready to deploy agentic AI securely, yet 83% plan to do it anyway. That gap between readiness and rollout is the threat surface. It's the threat surface.
Deloitte's 2026 AI report puts a finer point on it: just 20% of organizations have a mature AI governance model in place. Most are shipping agents before governance even exists to slow them down. And on the narrowest, best-known threat, prompt injection, only 34.7% of organizations have deployed a dedicated defense. The majority of enterprise AI deployments today are exposed to the attack class that's been publicly documented longest.
In December 2025, OWASP gave this problem a name and a shape. The OWASP Top 10 for Agentic Applications 2026 lists ten risks specific to autonomous agents: goal hijacking, tool misuse, identity abuse, agentic supply chain vulnerabilities, memory poisoning, cascading failures, and rogue agents, among others. Regulators are catching up too, though slower: the EU AI Act's high-risk obligations were pushed from August 2026 to December 2027, and Colorado's AI Act, as revised by SB 26-189, becomes enforceable January 1, 2027. None of that changes what's exploitable right now. This piece walks through the concrete bypass techniques attackers use against each layer of agent guardrails, and what each one reveals about where the defenses actually sit.
Why agents break traditional guardrail assumptions
Traditional guardrails were built for a simple shape: one input, one output, a human somewhere in the loop before anything happens. Agents don't work that way. They accumulate information across turns, cross organizational trust boundaries, call external tools and APIs, and can turn one bad instruction into a chain of real-world consequences before anyone notices.
Simon Willison's "lethal trifecta" explains why this is structural. Any agent that has access to private data (emails, documents, internal databases), is exposed to untrusted tokens (external emails, shared documents, scraped web content), and has some path to exfiltrate data (an external request, an image render, an API call, a generated link) is vulnerable regardless of how carefully its system prompt is worded. All three conditions together mean the agent doesn't need to be tricked in some exotic way. It just needs to do its job.
OWASP's Agentic Top 10 extends the older LLM Top 10 to cover risks that simply didn't exist before agents got tools and memory: tool-call hijacking, MCP server poisoning, multi-agent goal manipulation, identity confusion between agents, memory poisoning that persists across sessions, and privilege escalation through chained tool calls. Separately, the 2025 ATFAA threat model maps nine agentic threats across five domains, several of which target reasoning and trust boundaries that no prompt-level filter can reach, because the filter never sees the reasoning happen.
A 2025 agentic security survey (arXiv:2510.06445) makes the underlying mechanism explicit: wrapping a language model in an agentic framework significantly increases its vulnerability, because the safety alignment baked into the base model doesn't reliably carry over once that model is handed tools, memory, and autonomy. The guardrails built to inspect what goes into a model and what comes out of it are watching the wrong layer once the model can act.
The four control layers guardrails are meant to defend, and what each one protects
Guardrails for agents split into four functional layers, and each one has a distinct job and a distinct way of failing.
Input guardrails screen prompts and retrieved data before they reach the model: prompt injection detection, jailbreak filters, PII redaction, sanitization of content pulled in through retrieval-augmented generation. Output guardrails filter what the model sends back: toxicity and policy checks, secrets detection, grounding and citation verification. Topical and safety guardrails keep the model inside approved subject matter and enforce brand, legal, and ethical lines. Agent and tool guardrails govern what the system is allowed to actually do: tool allowlists, parameter validation, least-privilege scopes, human approval gates before high-impact actions fire.
Most production systems implement some of these layers and skip others. A deployment with strong output filtering but no input screening still lets a prompt injection walk straight through to the reasoning layer untouched. And an agent that produces perfectly clean text can still take an unsafe action, because the output layer and the action layer are two separate things watching two separate events. Gartner expects dedicated guardian agents, systems built specifically to police other agents, to hold 10 to 15% of the agentic AI market by 2030. That's a signal that agent-layer control is becoming its own product category, not an add-on feature.
Instructions in a system prompt are requests. A system prompt is probabilistic. The model can be argued out of following it. Real enforcement lives at the architectural layer, in callback functions, MCP gateways, and LLM gateways that intercept a request before it ever reaches a tool. At that layer, a block is structural. Nobody has to hope the model behaves.
Bypassing the model entirely: the CoreBreak infrastructure-layer attack
CoreBreak, disclosed by researchers Hedi Ingber and Aviyam Ivgi, co-founders of the security firm Stealth, at Black Hat USA 2026, found something uncomfortable: the tool-execution layers of Amazon Bedrock AgentCore, Google's Agent Development Kit, and Vercel's AI SDK harness packages could each be tricked into running a tool without a legitimate model turn happening.
If the model was never invoked, then system prompts, content filters, and refusal training are all beside the point, because there was no model decision for any of those controls to intervene in. The guardrails were skipped.
Four CVEs came out of the disclosure. The AWS Bedrock AgentCore harness is covered by CVE-2026-18830 (CVSS v4.0 8.6). CVE-2026-18236 (CVSS v4.0 9.3) covers the Google ADK flaw. CVE-2026-64650 and CVE-2026-64651 (CVSS v4.0 6.3 each) cover the affected Vercel AI SDK packages.
Monitoring for agentic systems has generally focused on model inputs and outputs, logging prompts, flagging odd completions, reviewing what gets passed to tools. When a tool fires without the model running at all, none of those artifacts exist. There's nothing to log, because the event the logging was designed to catch never happened. CoreBreak forces a specific conclusion: treating the model as the security perimeter is a mistake once the execution harness underneath the model is itself something an attacker can reach.
Making restricted content invisible to the filter that is supposed to catch it: tokenizer misalignment
Tokenizer misalignment works on a gap most teams never think to check: the guardrail and the underlying model don't necessarily tokenize text the same way. Content that a guardrail's tokenizer splits into harmless-looking fragments can reach the model intact, because the model's own tokenizer reassembles those fragments differently and reads the restricted content whole.
Research published as "Bypassing LLM Guardrails" at LLMSEC 2025 (arXiv:2504.11168) tested this against six well-known protection systems, including Microsoft's Azure Prompt Shield and Meta's Prompt Guard. In some cases the technique hit up to 100% evasion while keeping the adversarial content fully usable by the model on the other side.
This is a story about two separate systems, the guardrail and the model, disagreeing about what a string of text actually says, and that disagreement being the exploit. It's a story about two separate systems, the guardrail and the model, disagreeing about what a string of text actually says, and that disagreement being the exploit. No defender can assume that content passing a guardrail means the model downstream will interpret it the same way the guardrail just did.
Poisoning the tool-selection pipeline before the guardrail ever sees it: ToolHijacker
ToolHijacker attacks the part of the pipeline that decides which tools an agent even considers using. It's a no-box attack that needs no access to the model or the system it's targeting. It introduces malicious content into the retrieval pipeline the agent already trusts, shaping which tools the agent surfaces and acts on.
Against GPT-4o on the MetaTool benchmark, ToolHijacker hit an attack success rate up to 96.7%. And it didn't just slip past weak defenses. It bypassed both prevention-side defenses (StruQ, SecAlign) and detection-side defenses (known-answer detection, DataSentinel, perplexity detection, and perplexity windowed detection), the exact category of guardrail most teams reach for first.
The inspection point sits in the wrong place in the pipeline. Defenses focused on the agent's final tool call miss what happened earlier in the pipeline. By the time a guardrail looks at the decision, the poisoning already happened upstream of it. By the time the guardrail looks, the poisoning already happened upstream of it.
Exfiltrating data through the agent's own tool use: Imprompter and obfuscated adversarial prompts
Imprompter generates obfuscated adversarial prompts automatically, and demonstrated against production agents including Mistral LeChat and ChatGLM, coerces them into misusing their own tools to exfiltrate personal data. Measured extraction precision landed around 80%, and the technique transfers across model targets beyond the ones it was initially demonstrated against.
The obfuscation is the whole trick. It's tuned well enough that content filters and output guardrails never flag the instruction as malicious, because from the guardrail's vantage point, the agent's tool calls look like ordinary, permitted behavior. Nothing about the output text looks wrong. The data doesn't leave through what the model says. It leaves through what the model tells a tool to do, which is a channel most output guardrails were never built to watch.
EchoLeak (CVE-2025-32711, CVSS 9.3), discovered by Aim Security, showed exactly this pattern running in production. A single crafted email coerced Microsoft 365 Copilot into reaching into internal files and sending their contents to an attacker-controlled server, with no user interaction required. The attack chained four separate bypasses to get there: it evaded Microsoft's XPIA classifier, got around link redaction using reference-style Markdown, exploited images that auto-fetch, and abused a Microsoft Teams proxy that the content security policy happened to allow. Antivirus tools, firewalls, static scanning, none of it caught anything, because the exploit ran entirely in natural language. There was no code for those tools to inspect.
Attacking agents that can see: pixel-level injection in multimodal systems
WebInject targets agents that process a webpage as an image rather than as text. It optimizes pixel-level perturbations in a page's source that are imperceptible to a human eye but survive the (non-differentiable) process of turning that webpage into a screenshot the agent reads.
Across five multimodal web agents, WebInject achieved over 96% attack success using perturbations small enough that a human reviewer looking at the same screenshot sees nothing unusual. Standard input screening inspects text, the HTML, the visible copy, and the underlying page structure, so it has no way to catch this. The malicious instruction here doesn't live in any of that. It lives only in the visual representation the agent actually processes, a layer text-based filters were never built to read.
Different modalities and pipeline stages each open their own distinct injection paths, and no single text-based defense covers all of them at once. Every modality needs its own inspection layer, built for what that modality actually is. WebInject measuring results across five separate agents is a sign that multimodal web agents are common enough now that this isn't a hypothetical edge case anyone can defer.
Redirecting agent decision-making without injecting a prompt: control-flow hijacking
Control-flow hijacking skips the injected text altogether and goes after the reasoning path itself, the sequence of decisions an agent makes to get from a request to an action. Research into control-flow hijacking has found that even language models that resist ordinary prompt injection are still vulnerable to this class of attack, because resisting an injected instruction and resisting a redirected decision path are two different problems.
Rule-based defenses run into an unavoidable tradeoff here. Strict rules stop attacks but also stop the system from adapting to legitimate new situations, while flexible rules that allow adaptation can be talked around by a sufficiently clever request. No configuration removes that tradeoff. Someone always has to pick a side of it.
In multi-agent pipelines, the damage doesn't stay put. An attacker who hijacks one agent's goal can pass that corruption downstream, either by instructing the next agent directly or by poisoning shared memory that every agent in the pipeline reads from. The blast radius is the whole pipeline, extending beyond the single node that got compromised first.
Supply-chain research is now extending the same threat model further back, into the software agents are built from. BackdoorAgent (arXiv:2601.04566, 2026) offers a unified framework for backdoor attacks on LLM-based agents. SkillTrojan, presented at ICML 2026, targets backdoors specifically in skill-based agent systems. And in March 2026, a backdoor sat on PyPI inside a LiteLLM package for long enough to rack up close to 47,000 downloads before it was caught; that package serves as the language-model gateway for CrewAI, DSPy, and other orchestration frameworks, so one poisoned dependency reaches every agent built on top of it.
Coding agents carry the same exposure. GitHub Copilot's CVE-2025-53773 (CVSS 7.8) allowed remote code execution through a pull request description, and it's wormable across platforms. The Orca RoguePilot attack goes after Copilot Codespaces through a malicious GitHub issue, requiring nothing from the victim beyond opening a Codespace from that issue. Different product, same underlying failure: control-flow hijacking doesn't care whether the agent is answering customer emails or reviewing code.
What the bypass taxonomy reveals about where current defenses sit
Every technique above sorts into three failure modes, not ten.
Some guardrails inspect the wrong layer. CoreBreak bypasses model-level controls by never invoking the model. WebInject bypasses text-level controls by operating in pixel space, a layer those controls were never built to see. Others fail because the guardrail and the model disagree about what a given input actually means, which is the entire mechanism behind tokenizer misalignment. And in the remaining cases, ToolHijacker, Imprompter, control-flow hijacking, the malicious action is simply indistinguishable from normal, permitted behavior at the exact point where the guardrail is watching.
None of these are failures of effort. They're failures of architecture, the inspection point sitting in the wrong place, checking the wrong representation, or trusting that clean-looking output means clean-running underneath it. Guardrails built for a single model answering a single question were never going to hold against a system that plans, remembers, and acts across tools and sessions. The bypasses documented here aren't edge cases waiting to be patched one at a time. They're evidence that the whole layer needs rebuilding around what agents actually are, not around what language models used to be before anyone gave them hands.


