Est.

Least-Privilege Tool Access Controls for LLM Agents

Rethinking permission scoping and enforcement for AI agents that plan unpredictably across systems.

Staff Writer · · 11 min read
Cover illustration for “Least-Privilege Tool Access Controls for LLM Agents”
Runtime Controls · September 18, 2026 · 11 min read · 2,556 words

Least-privilege access for LLM agents isn't a checkbox on a security review. It requires rethinking how tool permissions get scoped, bound, and enforced while the agent is actually running, because agents plan and chain actions across systems in ways that static permission grants were never built to handle.

The principle itself is old news: give any actor only the access its current job needs, and nothing more. Banks have run on this logic for decades. What's new is the actor. A service account calls a fixed API in a fixed way, every time. An AI agent plans, picks tools, chains calls across systems, and changes course mid-task based on what an earlier step returned. Its next move can't be fully known at the moment someone grants it permission, because the agent's behavior is shaped by context that doesn't exist yet when the permission gets written down. One widely shared framing holds that the right way to think about this is to treat every agent as a full principal in its own right: a managed identity with a lifecycle, explicit roles, tightly scoped permissions, and a fixed manifest of tools it's allowed to touch. Security researcher Robert Saghafi's analysis points at the real weak spot: even systems that split planning from execution usually skip putting any real security check at the seam between the two. The planner decides to act. The system just assumes that decision was already authorized.

How badly the governance gap has already opened

Diagram: The Governance Gap: AI Agents vs. Security Readiness. Visualizes: Visualize the stark contrast between how widely AI agents are deployed versus how rarely organizations have controls in place to govern them.

Deployment has outrun governance, and the gap is wide enough to measure. Okta's AI at Work 2025 report found 91% of organizations are already running AI agents, but 44% have no governance framework at all, and only 10% have anything resembling a mature strategy for managing non-human identities. Netskope's AI Risk and Readiness Report 2026 found AI tools present in 73% of organizations, while real-time enforcement of governance policy had reached just 7%. Cisco's State of AI Security report put the number of organizations that feel ready to secure agentic AI at 29%, meaning 71% are running agents they can't properly watch, audit, or respond to when something breaks.

Visibility is a big part of the problem, and it compounds everything downstream. AvePoint's State of AI Report found 21.1% of organizations don't even know whether unsanctioned agent tools are running somewhere inside the business. Scoping permission for an agent nobody can see is not possible, full stop.

None of this is theoretical anymore. AvePoint's same report found 89.5% of organizations had at least one generative AI-related security breach in the prior 12 months, up from 75.1% the year before. IBM's Cost of a Data Breach Report found that among organizations that had an AI-related breach, 97% lacked proper AI access controls. That single number is probably the strongest argument in this entire piece: access control is the highest-leverage lever available, and most organizations aren't pulling it.

The scale only grows from here. IDC projects active AI agents in enterprises going from 28.6 million in 2025 to more than 2.2 billion by 2030. Every one of those agents has the same governance gap already built in, unless something changes structurally before they are deployed.

Excessive agency as a formally recognized risk category

Security researchers now have a name for this failure mode, and it's called excessive agency: an LLM-based system given the ability to call functions, reach systems, or trigger side effects beyond what its actual task requires. OWASP's Top 10 for LLM Applications formalized the category, and the risk isn't sitting still. It moved up significantly in the 2026 ranking compared to its position in 2025.

The field is maturing fast enough to split into its own discipline. OWASP's Top 10 for Agentic Applications 2026, released December 9, 2025, and built by more than 100 security researchers, treats agent-specific risk as distinct from chatbot risk, not a subcategory of it. Alongside the September 2, 2026 release of the 2026 OWASP Top 10 for LLM Applications, OWASP's GenAI Security Project also put out an Agent Control Standard and a new set of resources aimed specifically at securing agentic systems. A Dark Reading poll found 48% of respondents expect agentic AI to be the top attack vector for cybercriminals and nation-state actors by the end of 2026.

Excessive agency isn't just a label. It is the condition that makes the specific attacks below possible.

The attack paths that over-privileged agents open

Prompt injection is the front door. Vectra's research put attack success rates at 84% in agentic systems, and OWASP ranks it first on its Top 10 for LLM Applications (LLM01). The reason it's worse for agents than for chatbots is mechanical: an injected instruction in a chatbot produces a bad answer. An injected instruction in an agent can redirect a tool call, pull a credential, and chain that access into a downstream system the attacker never touched directly.

Security researcher Simon Willison named the underlying pattern the "lethal trifecta" in June 2025: an agent that has access to private data, processes untrusted content, and can talk to the outside world is exploitable by design, no matter how careful its prompt engineering is. Most deployed agents check all three boxes at once, because those three capabilities are usually what makes the agent useful.

MCP tool poisoning works a similar angle from a different door. An attacker hides instructions inside a tool's description metadata, and the agent reads that metadata as legitimate, trusted guidance about how the tool behaves. The MCPTox benchmark found many popular agents with attack success rates above 60%, topping out at 72%. Oddly, the most capable models often did worse here, not better, because better instruction-following just means more faithful compliance with whatever the malicious metadata told it to do.

EchoLeak (CVE-2025-32711, CVSS 9.3) shows what this looks like at full scale. Disclosed by Aim Security in June 2025, it was a zero-click flaw in Microsoft 365 Copilot: an attacker sends an email, and the agent exfiltrates confidential data with no user interaction. Microsoft patched it server-side in May 2025, and there's no confirmed evidence it was exploited in the wild before the fix. The chain behind it evaded Microsoft's XPIA classifier, got around link redaction using reference-style Markdown, exploited images that auto-fetch, and abused a Teams proxy allowed under the content security policy, stacking four separate bypasses into one privilege escalation across trust boundaries the system was supposed to enforce. An incident response team without logging tuned for this kind of system would struggle to reconstruct what happened.

Two more CVEs, GitHub Copilot's remote code execution flaw (CVE-2025-53773) and Claude Code's DNS exfiltration bug (CVE-2025-55284), confirm this isn't a lab exercise. These are production vulnerabilities with tracked vulnerability identifiers attached to real, deployed tools.

Scope creep tends to happen quietly, in small steps nobody flags for review. Microsoft's Security Blog describes the pattern: a team grants a broad "Reader" role because it looks safe enough, the workflow later expands to include write actions, and the team grants something wider than intended just to keep things moving, then never revisits it. Multiply that across every integration an agent has, and you get a combined-access problem that's worse than the sum of its parts: an agent with access to email, files, a ticketing system, and a code repo might look low-risk looking at any one connection in isolation, but the combination lets it correlate data across all four and take actions nobody explicitly signed off on as a package.

Static Permission Models and Agent Behavior

Traditional least-privilege design assumes a predictable actor. A service account calls a known API in a known way, a user opens a known file, and the permission set gets written once and reviewed on some fixed schedule. That model works fine when the actor's behavior doesn't change from run to run.

Agent behavior does change. Give the same agent the same goal twice, and it may plan a different sequence of tool calls each time, depending on intermediate results it hadn't seen before. A single static permission set built for that agent ends up either too wide, covering paths the agent never actually takes, or too narrow, blocking a legitimate path the agent needed this one time. This is an identity ambiguity problem: is the agent acting under its own identity, a delegated user's scope, or some blend of the two? That question determines who's accountable when something goes wrong, and it tends to be most acute during an incident, when logs show which tool got called but can't answer who authorized it, under what role, or whether the call was even inside the agent's intended scope.

Static roles simply don't carry task context. A "Reader" role granted for one workflow turns excessive the moment that workflow grows to include a write step, because the role itself has no idea what task the agent is currently doing. Multi-agent systems make the problem worse still: research on prompt injection has shown a malicious instruction can self-replicate across a mesh of interconnected agents almost like a virus, and a permission boundary drawn around a single agent does nothing to stop that kind of lateral spread.

What this points to is a hard requirement: permission has to be scoped to the task and the moment an agent is in.

Engineering task-scoped and time-bound tool access

AvePoint's framing captures the shift: least privilege for an AI agent means giving it access only to the specific data, systems, and actions its current task needs, only for as long as the task takes, and then pulling that access back automatically once it's done. That's a fundamentally different object than a standing permission set assigned once at setup.

In practice, this means each execution plan an agent runs should carry its own tools manifest, a fixed, declared list of exactly which tools this plan is allowed to call. Permission granted when a plan starts should expire when the plan ends, with no standing access carried over between separate invocations. Write access needs to be split apart from read access entirely, and granted only for the specific steps that actually require writing something.

Research has already produced working mechanisms for this. Progent, described by Shi and colleagues in April 2025, is a declarative policy language with per-tool constraints, guard conditions, and defined fallback behavior, enforced through wrappers around each tool call and a proxy layer sitting in the middle of every invocation. SEAgent, from Ji and colleagues in January 2026, builds mandatory access control on top of attribute-based access control, keeping a live graph of how information flows between agents, tools, databases, and user actions, with policy evaluated as logical expressions over paths through that graph. MiniScope, from Zhu and colleagues in December 2025, takes a different angle: it rebuilds the tree connecting OAuth scopes to API methods and solves for the smallest possible scope set that still lets a given execution plan complete, paired with runtime permission prompts.

A September 2025 paper by Jacob and colleagues pushes privilege separation down to the type level, restricting what agents can pass to each other, or between planning and execution stages, to a narrow, safe set of types, integers, enums, JSON that's validated against a schema, so that prompt injection has no channel to travel through. That's a mitigation that comes closer to a guarantee, because the injected text literally has nowhere to travel. It's closer to a guarantee, because the injected text literally has nowhere to travel.

The Plan-then-Execute architecture, described by Del Rosario, Krawiecka, and Schroeder de Witt in a September 2025 paper out of SAP and Oxford, turns the planner-executor split into an actual enforcement boundary rather than just an engineering convenience. Tool access gets scoped to individual executor steps instead of handed to the whole agent at once, and the executor only ever receives one step at a time, invoking just the tools that step needs. CrewAI already supports declarative tool scoping tied to agent roles. LangGraph uses stateful graphs to manage agent execution flow. AutoGen sandboxes code execution steps inside Docker by default. Adding a compliance check between planner and executor, before execution begins, matters most when the action in question can't be undone.

Enforcing least privilege at runtime, not just at configuration time

Scoping permission correctly at configuration time is necessary, but it isn't enough on its own. Agents run into untrusted content, unexpected tool output, and branching paths mid-task that no configuration written in advance could have fully anticipated.

Runtime enforcement means the permission check happens at every single tool call, not once at the start of a session. Each call gets evaluated against the task context that exists right then, rather than against one blanket token issued at login. Prompt Flow Integrity, described by Kim and colleagues in March 2025, splits the agent into trusted and untrusted parts, and forces every interaction with untrusted data through explicit, typed queries and proxies, with runtime checks that stop untrusted data from ever reaching a high-sensitivity tool call. Nested Least-Privilege Networks, from Rauba and colleagues in January 2026, go a level deeper, intervening directly in the model's weight matrices so that privilege level controls which internal computation paths are even reachable for a given request. Under that design, the model can't route around its own granted privilege level, because the paths that would let it simply aren't there.

For actions that can't be reversed, a human still belongs in the loop, not as a nice-to-have design pattern but as a runtime control in its own right; the verifier role in a Plan-then-Execute system can just as easily be a human expert as an automated check. And none of this holds up without logging built for it. Microsoft's Security Blog is direct about the requirement: logs need to capture who authorized an action, under what role, and whether it fell inside the agent's intended scope, not just which tool got called. Skip that, and post-incident investigation stalls no matter how tightly the original permissions were scoped.

Runtime enforcement is one layer in a stack that works alongside scoping access correctly at configuration time. Both layers are necessary, because each one catches failures that slip past the other.

Applying scoped access consistently across agent deployment platforms

The same discipline has to hold everywhere an agent runs, on every platform a security team manages.

Inside the Microsoft 365 environment, agents built through Copilot Studio, Microsoft Foundry, and SharePoint need lifecycle-managed identities and explicit role assignments through RBAC. Microsoft states that a managed identity paired with least-privilege RBAC is the floor, not some advanced configuration reserved for mature teams. EchoLeak is the clearest proof that platform-level controls alone don't hold: the exploit was possible specifically because the agent had broad access to email, files, and outside communication all at once, so reviewing combined access matters as much as reviewing each individual permission on its own.

For agents connected through MCP, tool descriptions themselves are part of the attack surface, since the agent reads that metadata as trusted, and tool poisoning exploits precisely that trust. Mitigating it means checking tool description content at the point of registration, restricting which MCP servers an agent is even allowed to connect to, and treating any third-party MCP server the same way a security team would treat unreviewed third-party code: with real scrutiny before it ever gets near production.

Sources

  1. Least Privilege for AI Agents: The Security Principle Your LLM Deployment Is Probably Violating | by Robert Saghafi | Medium
  2. AI Agent Least Privilege: A Practical Guide (2026) | AvePoint
  3. Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Implementations
  4. Least privilege for AI agents: Identity, access, and tool binding | Microsoft Security Blog
  5. arxiv.org
  6. The Granularity Mismatch in Agent Security: Argument-Level Provenance Solves Enforcement and Isolates the LLM Reasoning Bottleneck
  7. genai.owasp.org
  8. arxiv.org
Filed underRuntime Controls

More in Runtime Controls