Red Teaming vs Continuous Verification for Production AI Agents
Both red teaming and continuous verification are required, not alternatives.

Enterprise AI agents have stopped being chatbots that answer questions. They browse the web, run code, call APIs, write to memory, and chain tool calls together without anyone reviewing the intermediate steps, and that shift changes what "security" even means for the systems running them. Red teaming an agent before launch and verifying it continuously after launch aren't mutually exclusive options to choose between. Both are required, and most organizations are doing neither one well enough to close the gap between them. Treating the two as alternatives, one instead of the other, is the mistake driving most of the incidents showing up in the data right now.
Deployment has accelerated sharply across the enterprise, with embedded, task-specific agents moving from a niche capability to a mainstream expectation in the span of a year or two. Deployment has outrun everything else. Gravitee.io's April 2026 survey of 750 executives found that only 9.5% of organizations secure more than 81% of their deployed agents, with mean monitoring coverage sitting at 52%, meaning roughly half of production agents run with no real oversight at all. Industry research puts a finer point on it: only 14.4% of organizations send an agent to production with full security or IT sign-off, and 81% of security leaders say they feel pressured to ship before security work is finished. Eighty-eight percent of organizations running AI agents report a confirmed or suspected security incident in the past year, while only 6% of security budgets go toward agent security specifically. That gap between exposure and spend is the whole story in two numbers.
Part of the mismatch is a framing problem. Security teams still tend to think about AI risk as a model-level issue: jailbreaks, hallucinations, biased outputs. That's the wrong altitude to work at. An agent with valid credentials, live tool access, and a memory store operates in the world as an autonomous actor crossing systems, and the risk lives in the architecture, not the weights.
What makes agentic AI a qualitatively different attack surface
A conventional application has a boundary you can draw on a whiteboard, with inputs, a server, a database, and outputs. An agent's boundary moves. Every tool call it makes, every memory write, every message it passes to another agent expands what it can touch, and that expansion happens at runtime, not at deployment. You can't fully diagram an agent's attack surface before it runs, because the agent partly builds that surface as it goes.
Split the harm into two buckets. There are attacks on the agent, aimed at corrupting its goals, its reasoning, its outputs. And there are attacks through the agent, where the agent itself becomes the pivot point into a database, an internal API, or some other system it was trusted to reach. Most defensive thinking still fixates on the first bucket, and that's backwards. The second is where the real damage happens, because the agent already holds the credentials an attacker would otherwise have to steal.
Prompt injection makes the point directly. In a chat interface, a malicious instruction buried in a document or a webpage is an annoyance, maybe a bad answer. In an agentic system, that same instruction acts like remote code execution. If the agent reads a poisoned tool response or a compromised webpage and acts on what it read, and if that agent can call APIs, move data, or run commands, the injection turns into arbitrary action taken by a privileged process. Nobody typed a command. The agent just followed one it found lying around.
Traditional threat modeling wasn't built for this. STRIDE, the standard model for reasoning about spoofing, tampering, repudiation, information disclosure, denial of service, and elevation of privilege, assumes a system boundary that holds still long enough to analyze. An agent's boundary shifts with every tool it negotiates access to and every memory layer it populates during a session. Agent Goal Hijack, cataloged as ASI01 in the OWASP Agentic Top 10, has no clean STRIDE category at all: no spoofing occurs, nothing gets tampered with at rest, no privilege elevates in the classical sense. The agent just uses the capabilities it was already given, in service of a goal someone else planted. Component-by-component STRIDE analysis also misses attack paths that only emerge across multiple steps. EchoLeak, tracked as CVE-2025-32711 with a CVSS score of 9.3, looked fine when each component was checked on its own; it only turned dangerous once someone traced the full chain. Persistence through memory, an agent carrying a poisoned instruction forward across sessions, has no classical category whatsoever. The ATFAA framework treats temporal persistence as its own threat domain precisely because STRIDE never accounted for it.
Multi-agent systems compound the problem. A compromised sub-agent can pass manipulated output to an orchestrator, which treats that output as trusted input and executes high-privilege actions on the attacker's behalf, effectively escaping a low-trust sandbox into a high-trust execution context. He et al. (2025) found that Agent-in-the-Middle attacks exceeded a 70% success rate across most multi-agent configurations tested, and hit 98.5% in chain-structured agent communication specifically. Architecture isn't a neutral backdrop here, either. Research from Van hamme et al., accepted at DeMeSSAI 2026 and EuroS&P 2026 under the MATRA framework, shows that controls like network sandboxing and least-privilege access cut risk by shrinking the blast radius of a successful injection, regardless of whether the injection itself gets stopped. Architecture decides what's even possible, not just what the model happens to do.
Put together, this means the attack surface an agent exposes isn't fully knowable at the moment you deploy it. That single fact is why point-in-time testing, on its own, cannot close the whole gap.
What red teaming is actually designed to do
Red teaming an AI agent isn't the same exercise as red teaming a web application with a fixed codebase. The goal is finding the bugs that matter most, discovering how an agent's reasoning and tool-use patterns can be steered, coaxed, or tricked into outcomes an attacker wants.
Two approaches do this well, for different reasons. Manual red teaming leans on human creativity to surface attack vectors that automated tools tend to miss, particularly around goal hijacking, social engineering aimed at an orchestrator rather than a human, and multi-step chains that need the kind of intuition a script doesn't have. Automated fuzzing covers different ground entirely. Liu et al. (2025) built AgentFuzz and used it to find 34 zero-day taint-style vulnerabilities, including code injection, SQL injection, server-side template injection, and command injection, across 20 open-source agents, at 100% precision, resulting in 23 assigned CVEs. That's a strong result, and it's evidence that automated methods find real, exploitable flaws at a scale manual review simply can't match.
Formal frameworks now structure how this testing gets scoped. The OWASP Gen AI Red Teaming Guide, published in January 2025, lays out methodologies for both model-level and system-level vulnerability identification. MITRE ATLAS, at version 5.4.0, documents 84 adversary techniques and 42 case studies grounded in real incidents, with roughly 70% of its recommended mitigations mapping to security controls that already exist inside most organizations. NIST's ARIA pilot, version 0.1, ran about 51 red teamers through 508 testing sessions across seven submitted AI applications.
What a red team engagement produces, when done properly, is a set of concrete findings: a reproducible attack, demonstrated against the actual architecture, with a traceable impact path. That distinction matters more than it sounds like it should. A score is an opinion; a working exploit is a fact. Regulators are starting to treat it that way too. The EU AI Act imposes compliance obligations on high-risk systems by August 2026, and adversarial testing obligations under Article 55, aimed at general-purpose AI, are already in effect.
Where red teaming breaks down when applied alone
The first crack is non-determinism. LLM agents don't produce the same output twice under identical conditions, so a clean pass on one test run tells you less than it seems to. Research on repeated attack attempts found that running the same task across multiple attempts pushed average attack success rates from 57% up to 80%. A red team run proving an agent resists an attack once is not proof the agent resists it reliably.
The second crack is timing. Red teaming captures a snapshot: the attack surface as it existed during the engagement, and nothing about how that surface changes afterward. An agent that pulls in new external data, writes new entries to memory, or gets connected to a new tool after the engagement has acquired attack surface nobody tested. Agent-authored code makes this worse in a very literal sense. A coding agent can produce a new module in seconds, while a human threat model of that module takes hours to build, and in the gap between those two speeds, the agent can already have changed code, called tools, written memory, and updated dependencies, all before anyone checks the repository. Veracode's 2025 GenAI Code Security Report found that 45% of code samples failed security tests across more than 100 LLMs and 80 coding tasks, which gives some sense of how much unreviewed surface that window can generate.
Coverage is also bounded by whoever is doing the testing. Manual red teaming runs into scale limits and inconsistency between evaluators looking at the same system. NIST's AgentDojo-Inspect evaluation found that novel attacks reached an 81% task-hijack rate, compared to just 11% for prior baseline attacks. That gap says something uncomfortable: the space of possible attacks expands faster than any fixed test suite can track it. Rahman et al. (2025) demonstrated coordinated multi-agent attack chains reaching a 96.2% success rate against Claude 3.7 Sonnet, by spreading harmful intent across multiple turns rather than putting it in any single prompt. That's an attack class that emerges from how agents interact with each other over time. No red team fires that exact prompt, because there isn't one to fire.
There's a fix-verification gap too, and it's worth naming on its own. If a red team finds that a new safeguard breaks an existing workflow, that discovery usually surfaces at the next engagement, not the moment the safeguard goes live. Months can pass in between. Underneath all of this sits a question red teaming structurally cannot answer: is the agent still behaving as intended, right now, against this week's inputs, after whatever configuration change happened yesterday? That's a monitoring question, and it needs a different tool entirely.
What continuous verification adds and what it is not
Safety for a production agent is operational. You don't get it by running a test before launch and calling it done. You get it by constraining what the agent's environment allows, enforcing policy while the agent acts, watching its behavior as it runs, and keeping a way to intervene when something goes wrong.
Continuous verification does what red teaming structurally can't. It catches behavioral drift as the agent operates against actual production traffic, not synthetic test cases, which means it catches hallucination, goal drift, and configuration-sensitive vulnerabilities that only show up under specific real-world conditions nobody thought to script. It enforces guardrails at the moment of action, the tool call, the memory write, which is where prompt injection and context poisoning actually do their damage rather than at some abstract model boundary. And it builds an audit trail across sessions: evaluation probes embedded directly in live workflows generate verdicts that are structured, machine-readable, and usable as evidence for compliance review later.
AI agents don't just consume access the way a human user with a login does. They decide, session by session, how to use the access they've been given, and that moves the real control boundary away from login and entitlement issuance and toward verification at the moment of action, because that's where the danger actually lives. Simon Willison's "lethal trifecta," described in June 2025, names the condition precisely: an agent with access to private data, exposure to untrusted content, and the ability to communicate externally, all in the same session. Once those conditions converge, continuous monitoring stops being a nice-to-have and becomes the minimum viable control.
None of that makes continuous verification a substitute for red teaming. Watching a dashboard doesn't tell you whether a specific attack path actually succeeds against your agent's specific architecture, and detection without enforcement gives you visibility without control: the agent still takes the harmful action, and the team just finds out sooner rather than later. Human-in-the-loop review, too, doesn't scale against agents acting at machine speed across long-running workflows. Enforcement has to move at the agent's own speed, or it isn't enforcement at all.
Gravitee.io's 52% mean monitoring coverage figure is worth sitting with here, because coverage isn't the same thing as completeness. An agent that technically falls inside that monitored 52% can still have blind spots if the monitoring in place doesn't actually match its specific tool-call patterns.
How the two approaches fit together across an agent's lifecycle
Red teaming and continuous verification sit at different points in an agent's life and answer different questions. They aren't competing for the same budget line or the same practitioner's time. Treating them as substitutes for one another is where most of the gap described above comes from.
Red teaming owns the pre-deployment window and the change-event window. Before an agent goes live, it proves which specific attack paths succeed or fail against the actual architecture and sets a known-good baseline. Before any significant change, a new tool integration, a new memory backend, a reconfigured orchestrator, a model swap, a re-engagement scoped to just that change catches drift before it piles up silently in the background. What matters about the output is that it's a proven attack with a demonstrated blast radius rather than a score. Scores can't be reproduced. Attacks can.
Continuous verification owns everything in between. It watches for behavioral drift against real production traffic in the gaps between red team cycles. It enforces policy right at the point of action, tool calls, memory writes, data leaving the system, so a known attack class fails even when the exact payload wasn't in any test suite. And it generates the audit trail that makes each decision defensible to whoever has to approve a release or answer to a regulator later.
The two are supposed to feed each other. A finding from a red team engagement should become a standing detection rule inside continuous verification, not a closed ticket that gets filed and forgotten. Research on tool poisoning gives a sharp reason why this matters: tool poisoning attacks slip past output-based safety checks in 95% of cases, which means monitoring that only inspects outputs is nearly useless against this attack class. Enforcement has to happen at the tool-call layer itself. CSA's MAESTRO framework, published in February 2025, adds another layer to the argument: the most dangerous attack paths tend to start at the foundation model layer and cascade all the way down through deployment and infrastructure. No single red team engagement and no single monitoring dashboard covers every layer that framework describes. A combined posture is the only one that can.
What a coherent security posture for production agents looks like in practice
Start from the threat model, not from whichever tool is easiest to buy. MATRA's asset-based impact assessment approach connects specific confidentiality, integrity, and availability impact scenarios to attacker objectives and to the vectors a particular architecture actually exposes, before anyone decides what to test or what to watch.
Red team scope should follow the agent's real tool surface and trust boundaries, not a generic test plan written for language models in the abstract. An agent with write access to a production database and an agent that only reads public documentation are not the same threat model, even if they run on the same underlying model. Testing them identically wastes the exercise.
Continuous verification should enforce at the tool-call and memory-write layer specifically, given that output-based checks alone miss the large majority of tool poisoning attempts. Detection without enforcement is visibility, and the two get confused constantly, usually to the detriment of whoever assumed a dashboard alert counted as a control.
The findings from every red team engagement belong inside the detection rules continuous verification runs day to day. A proven attack that doesn't become a standing test is a wasted engagement, no matter how good the report looked. And the re-engagement cadence should track architectural change, new tools, new memory backends, new orchestrator logic, rather than sitting on a fixed calendar that has nothing to do with when the attack surface actually moved.
This is essential rather than optional polish. With 88% of organizations running agents having already logged a security incident, and only 6% of security budgets pointed at the problem, the gap between deployment speed and verified security posture isn't shrinking on its own. It closes when red teaming and continuous verification work as one system, each covering exactly the window the other can't.


