Exfiltration Prevention Controls for Tool-Calling Agents
Four control layers stop agent data theft at input, tool use, output, and memory.

Data exfiltration from a tool-calling agent doesn't happen at a single point of failure. It happens across a chain: when untrusted content enters the agent's context, when the agent picks a tool and fills in arguments, when it sends a result somewhere, and when it writes something to memory for later use. Four separate control layers are needed to stop it, one for each link in that chain, and skipping any single layer leaves a working exfiltration channel open no matter how strong the other three are. Most teams pour their budget into input filtering, the layer attackers have already learned to route around, and leave the tool and memory layers running on trust.
Traditional security tooling assumes a human starts a discrete action with a known scope and a clear endpoint. Agents break all three assumptions at once. A web application firewall doesn't parse an agent's reasoning chain. A data loss prevention tool doesn't look inside an LLM's context window. A cloud access security broker can't tie a tool call back to a human identity when the call comes from an autonomous process running under a service account. These tools were built to watch people, and an agent moving through five systems in under a second is one process carrying the combined authority of everything it touches.
Agents chain steps together on their own, and each step can look completely fine in isolation. A bulk file read, followed by an API write, followed by an outbound upload, reads as three unremarkable actions to any system watching one at a time. Only the sequence gives it away. Agents also read documents, web pages, API responses, and database records as part of how they think, so an attacker's instructions can ride in on that same channel as trusted data, no different in form from a legitimate paragraph of text. Once that instruction lands in the context window, it doesn't stay contained. State carries from one tool call to the next across a session, so a single poisoned document can corrupt everything that follows it.
Speed compounds the problem. According to TrueFoundry's buyer's guide for enterprise AI security, a full exfiltration chain, from injected instruction to data leaving the building, can run its course in seconds, well before a human reviewer notices anything happened. The identities involved often aren't human at all: database credentials, OAuth tokens, API keys, tied to a service account with no individual employee behind it in the identity management system. Obsidian Security's analysis makes the point directly. An agent with access to several enterprise systems at once doesn't put one person's data at risk; it exposes the combined authority of every permission it holds, across every system it touches, at the same time.
The cost of getting this wrong is already measurable. IBM's Cost of a Data Breach Report put the average cost of a shadow AI breach at $4.63 million, roughly $670,000 above a standard breach. That gap isn't just about scale. It reflects a structurally different kind of exposure, one that standard breach response wasn't built to find or contain.
The data path through a tool-calling agent and where exfiltration can branch off
TrueFoundry's enterprise buyer's guide breaks the agent stack into five layers where threats show up: identity (credentials, tokens, session context), model and prompt (the LLM call itself and what's sitting in the context window), tool and MCP (server connections, tool invocations, outbound API calls), runtime behavior (multi-step execution and state that persists across sessions), and compliance and audit (the trail left behind for governance and regulators). Exfiltration risk touches all five, but it concentrates hardest at the tool layer. That's where reasoning turns into action, and action is the only step that can actually move data anywhere.
A tool call looks simple from the outside, but it's a bundle of separate parts, each of which can be compromised on its own: the tool selected, its schema, the argument values passed in, the result that comes back, and whoever consumes that result downstream. Treating this as a provenance problem is the right frame. A completely legitimate, fully authorized tool turns dangerous the moment its parameters get filled in from untrusted or malformed input. The tool didn't do anything wrong. The data feeding it did, and that distinction is the whole ballgame for where defenses need to sit.
OWASP's Agentic AI Top 10, published December 9, 2025 as the ASI 2026 list, names the relevant risk categories directly: goal hijacking (ASI01), tool misuse (ASI02), identity and privilege abuse (ASI03), memory poisoning (ASI06), and insecure inter-agent communication (ASI07). Research cataloged under Owner-Harm category C6 in arXiv:2604.18658 describes the exfiltration pattern in plain terms: the agent uses an authorized tool, such as email, a webhook, or a file write, as a covert channel to smuggle data to an endpoint the attacker controls. Nothing about the tool call itself looks unusual. The destination is the tell, and destinations are what most tool-layer monitoring ignores.
Multi-agent setups make this worse fast. Analysis cited by Cycode found that connecting five MCP servers to a single agent let one compromised server hit a 78.3% attack success rate, and that compromise cascaded into the operations of the other connected servers 72.4% of the time. A single weak link in a tool network doesn't stay contained to that link. Treating each server as independently trustworthy is the mistake that lets it spread.
Input-layer controls: stopping injected instructions before they reach the reasoning pipeline
Prompt injection isn't a bug waiting for a patch. Sysdig frames it correctly: large language models can't tell instructions apart from data, because both arrive as the same thing, natural-language tokens in a sequence. That's a property of how the architecture works, not a flaw somebody forgot to fix, and no amount of prompt engineering changes it.
Direct injection is the easier case: a user types something designed to override the system prompt. Indirect injection is harder, because the attacker never talks to the model directly. Instead, the malicious instruction sits inside a document, a webpage, an email, a retrieval result, or an MCP tool's output, and the agent picks it up as part of doing its job normally. This touches nearly every kind of agent in production: RAG systems, browsing agents, MCP tool servers, and email-summarizing assistants all read attacker-controllable text as a routine part of their operation. It's the default mode of operation for anything that reads text it didn't write.
Some defenses attack this at the structural level. Delimiter and tagging schemes mark untrusted content as untrusted the moment it enters the context window. Others borrow directly from a decades-old fix for SQL injection: separate the instruction channel from the data channel before anything gets interpreted, the same logic behind prepared statements. arXiv:2605.18991, "Agent Security is a Systems Problem," proposes taint tracking on any byte that originated from user or external input, flagging its provenance before it ever reaches the reasoning loop, similar in spirit to how NX bits or parameterized queries work in traditional software security.
None of this closes the door completely, and teams that treat input filtering as the finish line are the ones who get burned. The 2025 paper "The Attacker Moves Second" (Nasr et al., arXiv:2510.09023) found that adaptive attackers who know what defense they're facing can route around nearly all published input-layer protections. Five carefully crafted documents, planted through RAG poisoning, manipulated model responses 90% of the time in that study. The attack surface isn't limited to text, either. CrossInject, presented at ACM MM 2025, achieved at least a 30.1% improvement in attack success rate over prior adversarial image methods, and image inputs remain a gap most teams haven't even started closing.
Input controls narrow the attack surface. They don't seal it, and no vendor claiming otherwise should be trusted. Anything that gets through still has to pass through the tool invocation layer before it can do damage. That is where the real enforcement work has to happen.
Tool invocation controls: enforcing what the agent is allowed to call and with what arguments
Tool misuse (OWASP's ASI02) and excessive agency (LLM06:2025 in OWASP's LLM Top 10) both live at this layer. A compromised model doesn't need an unauthorized tool to do harm; it just needs to chain authorized tools into a sequence nobody approved. It just needs to chain authorized tools into a sequence nobody approved. Each individual call passes muster. The chain doesn't, and judging calls one at a time is how that chain slips through undetected.
The starting fix is least-privilege scoping, applied at the level of individual tool grants rather than whole systems. An agent's role should define exactly which tools it can call, and read access and write or transmit access need to be separated cleanly: an agent that reads a CRM record has no legitimate reason to also hold a send-email tool. Grants should be time-bound and scoped to a session wherever possible, rather than standing indefinitely, because a permission that never expires is a permission an attacker eventually finds.
That still leaves a granularity problem. Most existing defenses mediate trust at the level of the whole tool call, an allow-or-deny decision with no room in between, too blunt to catch what happens inside the call. What's actually needed is inspection at the argument level: before execution, trace where a given value came from. Did that destination URL come from the trusted system prompt, or did the agent lift it from an untrusted document it retrieved five steps earlier? Google's CaMeL framework (Debenedetti et al., 2025) answers this by separating control flow from data flow explicitly, acting as a reference monitor that attaches security policy to individual data flows and checks for violations before anything runs. A 2026 extension called CaMeLoT (arXiv:2609.18674) goes further, adding a static check on the agent's entire plan before a single tool fires, which moves part of the enforcement earlier than runtime.
MCP introduces its own set of risks on top of this. A malicious MCP server can advertise a tool whose description is written to manipulate which tool the agent picks, a technique known as tool poisoning. Supply chain compromise is a live threat too: TrueFoundry documented the ClawHavoc campaign, in which attackers published more than a thousand malicious skills to ClawHub, OpenClaw's skill marketplace, deploying credential stealers onto enterprise developer machines through what looked like ordinary marketplace listings. Countermeasures here include cryptographic attestation of MCP server identity, an allowlist of approved tool schemas, and alerting whenever a registered tool's description changes without warning.
Multi-agent pipelines compound all of this. "Prompt Infection" (Lee & Tiwari, arXiv:2410.07283) documented injected instructions spreading from one agent to the next across a pipeline, and a coordinated 2025 attack reached a 96.2% success rate against Claude 3.7 Sonnet. Tool invocation controls can't just sit at the entry point of a multi-agent system. They need to apply at every node in the graph, because an injection that clears the first agent can still spread to every agent downstream of it.
Output routing controls: inspecting and constraining what the agent is allowed to transmit
Even a well-scoped, argument-checked tool call can still leak data. The exfiltration pattern in arXiv:2604.18658 doesn't require breaking any rule about which tool gets used: the agent sends an email, hits a webhook, or writes a file, all actions it's fully authorized to take, and simply routes the output to an address the attacker controls. The invocation is legitimate. The destination isn't, and tool-layer scoping can't see that gap.
Catching this means inspecting outbound arguments that represent destinations, URLs, email addresses, file paths, and checking them against an approved-destination allowlist before the call executes. It also means running sensitive-data detection on whatever the agent is about to send: PII, credentials, financial records, internal identifiers. Detection without an enforcement action does close to nothing here. Redaction and blocking have to be live options that act in the moment, before the data leaves the building.
There's a real gap between how legacy tooling handles this and what agent behavior actually requires. Nightfall AI's 2026 platform documentation puts precision from ML- and LLM-based detection at 95% out of the box, against a 5 to 25% accuracy baseline for legacy DLP tools built around human behavior patterns. Legacy DLP wasn't built to watch a process that reads a thousand files and sends an email in the space of four seconds, so it either misses the pattern or drowns real events in noise. Nightfall's own guidance makes the point directly: agents chain calls and move data without waiting for human approval, so blocking has to happen in real time. Alerting after the fact arrives too late to matter.
EchoLeak (CVE-2025-32711) is the clearest proof this isn't theoretical. It demonstrated a zero-click exfiltration path, no user interaction required at any stage. An output control built on the assumption that a human confirms each transmission before it goes out simply doesn't apply to an attack shaped like that, and most attacks now are shaped exactly like that.
Memory access controls: treating the persistent memory store as a distinct exfiltration vector
Everything covered so far operates within a single session. Memory breaks that boundary. Long-term memory systems let an agent store, retrieve, and reuse information across sessions, and an attacker who successfully writes something to memory doesn't just compromise one conversation. That entry sits there, waiting to shape whatever the agent does next, tomorrow or next week, with no fresh injection required.
The provenance gap here is worse than at the input layer. Research in arXiv:2606.04990 notes that most current memory systems offer little way to trace where a stored item came from, whether it's still accurate, whether it conflicts with something learned more recently, or how it ends up shaping a later decision. Two 2026 papers demonstrate what that gap allows in practice. Trojan Hippo (Das et al., arXiv:2605.01970) shows persistent memory functioning as its own exfiltration channel, separate from anything happening in the live session. Hijacking Agent Memory (Wang et al., arXiv:2605.29960) shows the poisoning doesn't need anything dramatic; ordinary-looking conversational turns are enough to plant a corrupted entry. Earlier documented cases of AutoGPT memory poisoning, cataloged in arXiv:2604.18658, showed injected memory entries causing the agent to act on an attacker's behalf across sessions that came afterward, with no new injection needed each time.
Closing this gap takes a handful of specific controls, and none of them work alone. Every memory write needs a provenance tag: source, session ID, and a trust level attached to the content that produced it. Before a stored memory item gets used as an argument in a tool call, its provenance tag needs to clear whatever trust threshold that action requires. Memory items sourced from untrusted external content shouldn't persist indefinitely; they need a TTL or a re-validation step before they're trusted again. And when a new memory entry contradicts something already stored at a higher trust level, the system should flag the conflict rather than silently letting the new entry win.
OWASP's Agentic AI Top 10 names memory poisoning as its own category, ASI06, and that placement signals where this is headed: compliance frameworks are going to start asking for documented controls at this layer specifically, not folded into a general data-handling policy.
Composing the four control layers into a working enforcement architecture
None of these four layers substitutes for the others, and treating any single one as sufficient is the most common mistake in this field. Input controls catch what they catch and let the rest through. Tool invocation controls stop unauthorized chains but don't inspect where a legitimate transmission is headed. Output routing catches destination-based exfiltration but has no visibility into what's sitting in memory from three sessions ago. Memory controls address persistence but assume the session-level layers already did their job upstream. A working architecture runs all four at once, each one catching what the others structurally can't.
The framing in arXiv:2605.18991, authored by researchers across Google, UCSD, UW-Madison, EmbraceTheRed, Gray Swan AI, Cornell, and other institutions in May 2026, states the underlying principle: the model driving the agent has to be treated as an untrusted component. Security has to be enforced by the system built around the model, not by trying to make the model itself more careful. Every effort to improve model robustness helps at the margins, but none of it removes the need for external enforcement, because the model stays the one part of the system an attacker can talk to directly.
That points to a small set of systems-security principles, borrowed from decades of practice in traditional software security, that map directly onto the four layers described here. Minimum privilege limits what an agent can even attempt, cutting off entire categories of tool misuse before a single call happens. Explicit mediation, checking every tool call, every argument, and every memory read against policy rather than trusting the model's own judgment about what's safe, is what turns good behavior from a hope into an enforceable guarantee. Provenance tracking, threaded through input tagging, argument tracing, output inspection, and memory tagging alike, is the connective tissue that lets every other control function. None of them work without knowing where a piece of data came from.
Agents will keep getting faster and more autonomous, and the gap between what they can do in a single session and what a human can review in that same window will keep widening. The four-layer architecture described here is the current shape of a containment strategy built for a threat surface that changes as fast as the agents running against it.
Sources
- Best AI Agent Security & MCP Security Platforms for Data Exfiltration Prevention in 2026 | Nightfall AI
- Enterprise AI Agent Security Solutions: The Complete Buyer's Guide (2026)
- Agent Security is a Systems Problem
- Securing AI agents: the defining cybersecurity challenge of 2026
- Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration
- Prompt Injection Attacks on AI Agents: How to Detect and Prevent Them
- arxiv.org
- The Comprehensive Guide to Prompt Injection Attacks in 2026 | Sysdig


