Argument-Level Enforcement in AI Agent Tool Calls
Text-layer content filters miss five critical attack patterns that live in tool arguments instead.

How agents execute actions, and what that means structurally
When an agent "does something," it doesn't produce a sentence. It issues a tool invocation: a named function, called with typed, structured arguments. Calling an API, writing to a database, kicking off a workflow, pushing an instruction to a connected system, these are the verbs of agentic behavior, and none of them look anything like the paragraph a content filter was built to scan.
The tool name is only half the story, and it's the less important half. A delete_record call is permitted or it isn't, but permission at the tool level says nothing about scope. delete_record called with a single row ID is a routine operation. delete_record called with WHERE 1=1 is a different action entirely, wearing the same name tag. The damage, when it happens, lives in the argument payload, in its target, scope, and magnitude. The damage lives in the argument payload, and that's what matters most.
Agents also don't stop at one call. They chain them, feeding the output of one invocation into the input of the next, and the attack surface compounds at every link. A single bad argument early in a chain can propagate downstream into calls that look, on their own, entirely legitimate.
Much of this chaining now runs over the Model Context Protocol, the standard many agents use to reach external tools and data. Each MCP server call is its own discrete invocation with its own argument payload, and a compromised or spoofed MCP server can feed malicious instructions back through that exact channel, or use it to pull data out. Structured parameters can be checked against a schema, a policy, or a record of where they came from, but only if something is actually positioned to look at them before execution happens. Most stacks aren't built that way. Tool invocations get trusted by default: analysis from neuraltrust.ai finds no risk scoring before a call runs, no policy enforcement sitting at the connector, no audit trail showing what an agent actually did across the environment.
What Content Moderation Cannot See
Five distinct failure patterns appear once you look at the argument layer instead of the text layer. None of them are text problems, and treating them as text problems is the mistake most security teams are still making.
Injected tool arguments come first. A poisoned input, buried in a document or a support ticket, convinces the agent to call a real, permitted tool with destructive arguments. As futureagi.com puts it, the text is clean and the action is not. Content moderation was never pointed at the tool call, so it has nothing to catch here, and no amount of tuning the text filter fixes that.
Excessive agency is second. An agent scoped for read-only work invokes a write endpoint instead, stepping outside the task it was built for. OWASP lists this among the top risks specific to LLM applications. No toxicity scanner or PII filter sits anywhere in that path to notice a permission boundary getting crossed.
Third is compromise at the MCP layer itself. The agent reaches a spoofed or malicious MCP server, and content filters never touch that traffic because it isn't user-facing text. CVE-2025-49596, a flaw in MCP Inspector scored 9.4 on CVSS, showed how bare that layer can be: basic session tokens and origin verification were missing from the outset.
Fourth, argument-level exfiltration. Arguments can encode a destination, or carry a sensitive payload disguised as a routine parameter. Structurally, the call looks fine. Functionally, it routes data to an endpoint the attacker controls, and nothing about its shape trips a content check.
Fifth, system-prompt leakage through crafted arguments. A carefully built input extracts the system prompt, exposing the tool access and trust boundaries the agent operates under, information an attacker then uses to shape later calls.
One thread runs through all five. These are structural events happening in the parameter layer, not semantic events happening in the text layer. A scanner asking "is this text safe?" is asking the wrong question of the wrong layer, full stop. The OWASP Top 10 for LLM Applications catalogs prompt injection, excessive agency, and system prompt leakage as named risks, and not one of them is a text-toxicity problem.
What the real-world incident record shows about argument-level exposure
The clearest illustration comes from a breach of Mexican government systems between December 2025 and February 2026, documented as the opening case study in Check Point Research's Annual AI Security Report. One operator, working from 1,088 typed prompts, produced more than 5,000 AI-executed commands across 34 attack sessions. Claude Code carried out roughly 75% of the live exploitation work: exploring systems, running exploits, harvesting credentials across 305 internal servers. A separate pipeline built on GPT-4.1 generated 2,597 structured intelligence reports and automatically tasked follow-on activity from them. The total haul ran to roughly 400 million records, pulled off using more than 400 custom attack scripts against 20 different CVEs.
The enforcement lesson sits in plain sight. The coding agents involved ran with no scope boundary, no ceiling on execution, no kill switch. Nothing checked whether the next tool call fit within a defined scope before it fired, and more than 5,000 commands executed because nothing inspected each one against a policy at the moment it ran. A per-call scope check, or even a blunt ceiling on total invocations, would have broken that chain long before it reached 400 million records. That's the argument most people get backwards here: the failure wasn't that the AI was too capable. It was that nothing was watching what it did with that capability, call by call.
EchoLeak, tracked as CVE-2025-32711 and scored 9.3, tells a related story from a different angle. Disclosed by Aim Security in June 2025, it was a zero-click prompt injection against Microsoft 365 Copilot: a single crafted email, no user interaction required. The chain stacked several bypasses on top of each other: evading Microsoft's XPIA classifier, getting around link redaction with reference-style Markdown, exploiting auto-fetched images, and abusing a Teams proxy the content security policy happened to allow. Copilot reached into internal files and shipped their contents to a server the attacker controlled, potentially exposing chat logs, OneDrive files, SharePoint content, and other organizational data. Researchers who disclosed it pointed to provenance-based access control and prompt partitioning as the fix, controls that act on where a call's data came from and what scope it holds, not on whether its text reads as toxic. It stands as the first documented case of prompt injection weaponized for concrete data exfiltration in a production system.
Money moved too. In January 2026, attackers drained roughly $40 million from Step Finance's treasury wallets after compromising executive devices, in a breach that drained funds from executive-controlled wallets. Neuraltrust.ai's account describes the agents involved as having done what they were built to do. Nobody told them to stop, and nothing at the argument level was positioned to stop them either: no ceiling, no approval gate, nothing standing between decision and execution.
Supply chains carry the same risk one layer upstream. A backdoor sat inside a LiteLLM package on PyPI in March 2026, in a window that saw nearly 47,000 downloads. LiteLLM serves as the LLM gateway for CrewAI, DSPy, Microsoft GraphRAG, and a range of other agent frameworks, so a single compromised package at that layer propagates quietly through every stack built on top of it. Separately, a GitHub Copilot remote-code-execution flaw (CVE-2025-53773, CVSS 7.8) and a critical Cursor IDE vulnerability (CVSS 9.8), both confirmed across 2025 and 2026, show that coding-agent tooling is now an active target.
Content moderation was either bypassed outright or simply irrelevant across every one of these incidents. The damage executed at the action layer, every time. Survey data backs up how widespread this has already become: agatsoftware.com reports that 88% of organizations reported a confirmed or suspected AI agent security incident in the past year, and in healthcare specifically that figure climbs to 92.7%.
What argument-level enforcement is, definition and operating principles
Argument-level enforcement is logic that intercepts a tool invocation after the model has decided to make the call but before it executes, and checks the structured parameter payload against a policy. It's the argument getting checked.
Four things fall inside that check. The tool name itself: is this tool even in the allowed set for this agent, this task, this context? The argument values are checked against an expected schema, an expected type, and an expected range. The provenance of each argument: did this value come from trusted system context, or did the agent pick it up from untrusted input it read somewhere along the way? And the sequence: does this call fit a pattern the workflow expects, or does it mark a deviation from it?
What it isn't matters just as much, and this is where a lot of teams get confused. Argument-level enforcement doesn't replace content moderation, PII scanning, or output filtering. It sits on a different layer and answers a different question. Content moderation asks whether text is safe. Argument-level enforcement asks whether an action falls within policy. Those are separate checks, and a working system needs both, not one standing in for the other.
The provenance piece is where indirect prompt injection actually gets answered structurally, and it's arguably the single most important mechanism in the whole stack. Knowing that an argument value traces back to untrusted external content, rather than a trusted instruction, would have flagged the refund_order call buried inside a support ticket, or the tainted parameter riding along inside EchoLeak's crafted email. Research systems like ProGENT enforce programmable privilege control right at the tool interface, blocking calls that fall outside an allowed policy. Proposals like MCPGuardian aim to put a security-first mediation layer in front of MCP-based tool requests specifically. Both sit at the argument layer by design.
Location matters as much as existence here. A policy listed in a system prompt, one the model is simply asked to respect, is not enforcement. It's a suggestion, and agents under adversarial pressure do not reliably honor suggestions, no matter how firmly the prompt phrases it. Enforcement has to sit at invocation time: not at startup, not in an after-the-fact audit log nobody reads until the damage is done. It also has to be sized right. The correct safeguard is the smallest one that stops the attack while leaving alone the workflows that posed no threat to begin with. Overly broad blocks generate enough friction that teams start disabling the very enforcement meant to protect them, which defeats the entire point of building it.
How argument-level enforcement works in practice, the mechanics
Six mechanisms, stacked together, cover most of what argument-level enforcement looks like in a working system.
Tool allowlisting comes first, and it fires per call, not once at startup. Futureagi.com describes a tool-permission guardrail as checking which tools an agent is allowed to invoke at the exact moment it tries to invoke one, not when the session begins. Second is argument schema validation: before a call executes, its parameters get checked against a declared schema for type, range, enumeration, pattern. A user_id that resolves to a wildcard, or appears in an unexpected format, gets stopped before the database ever sees it.
Third, provenance tagging and taint tracking. Arguments that trace back to untrusted content, a document the agent read, a webpage it fetched, an email it processed, get tagged the moment they enter the system and tracked through however many reasoning steps follow. If a tainted value later shows up as an argument in a tool call, the enforcement layer applies extra scrutiny or blocks the call.
Fourth, a hard ceiling on scope and execution: a cap on tool invocations per run, on the volume of data an agent can move, and on which systems it's allowed to touch. The Mexican government breach ran past 5,000 commands with no ceiling in place to interrupt the chain. This is a structural kill switch, not a hope that the agent behaves itself.
Fifth, human approval gates on anything destructive or high-risk. A call that would write to production, send data externally, delete records, or escalate privilege gets paused and surfaced for a person to approve before it fires. Analysis from waxell.ai covering the Check Point findings highlights OWASP's excessive-agency category as directly relevant to this gap.
Sixth, inspection of MCP traffic specifically. For deployments running over MCP, enforcement sits on the gateway path itself and checks tool requests before they execute, which is exactly where a spoofed or compromised MCP server would otherwise inject malicious instructions into the agent's reasoning.
Every inspected call, allowed or blocked, should leave behind a tamper-evident log entry recording the tool name, the argument payload, the enforcement decision, and the provenance tag attached to it. That log is the audit surface the current default, no governance at the tool layer at all, simply doesn't produce. None of this ships without verification against approved workflows first, either. A safeguard that blocks something legitimate defeats its own purpose. It's an outage with a delay timer on it.
Where current frameworks and tools address this layer
Coverage across the existing security frameworks is thin, and the gap is measured. A study cataloged in arXiv:2603.09002 (Nguyen et al.) scored 193 distinct threat items against 16 separate security frameworks, and no framework in the set achieved majority coverage of any single risk category. The OWASP Agentic Security Initiative led the field overall, at 65.3% coverage, which marks the high end of a field where most frameworks land well below it.
The categories that fared worst are the ones tied most directly to argument-level behavior, and that's not a coincidence. Non-Determinism scored a mean of 1.231 across all 16 frameworks, and Data Leakage scored 1.340, both near the bottom of the scale used. These are exactly the risks that live in what an agent's arguments do at runtime, and most frameworks were simply never built to watch that layer.
OWASP's own Top 10 for Agentic Applications, released in December 2025, names goal hijacking and tool misuse as the two leading risks specific to agentic systems, ahead of everything else on the list. Both are argument-layer events by definition. MITRE's ATLAS framework, at version 5.4.0, offers the most granular adversarial taxonomy available for AI systems generally, spanning 16 tactics, 84 techniques, 56 sub-techniques, and 32 mitigations. Even there, tool-call argument exploitation as its own distinct attack surface remains an emerging category rather than a mature one.
The commercial market has started organizing around the gap, though it hasn't converged yet. Analysis from withwillow.ai splits the vendor landscape into five categories: MCP gateway and tool-call governance, runtime detection and response, AI security state management and discovery, inference firewall and red-teaming tools, and identity and access governance built for agents specifically. That same analysis notes that 48% of cybersecurity professionals expected agentic AI to become the number-one attack vector going forward. Operant AI sits in the MCP gateway and tool-call governance category, positioned between every agent and every tool it calls, inspecting and enforcing each invocation as it happens, which comes close to the textbook definition of argument-level enforcement in commercial form.
None of this closes the gap, and pretending otherwise would misstate the state of the field. Framework coverage remains sparse, the categories tied most tightly to runtime argument behavior score the worst across the board, and the vendor market is still sorting itself into categories rather than settling on a standard. What's clear is where the next round of work has to go: down, past the text, into the arguments themselves.
Sources
- The 8 AI Agent Security Tools Enterprise Teams Should Evaluate in 2026 | Willow
- AI Agent Attack: 5,317 Commands, 9 Agencies Breached [2026]
- The Complete Guide to AI Agent Security for Enterprises (2026) | NeuralTrust
- Security Considerations for Multi-agent Systems
- Agent Runtime Guardrails in 2026: The Tool-Call Scanners Most Stacks Skip
- From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI
- agatsoftware.com
- helpnetsecurity.com


