Privilege Escalation Paths in LLM Agent Tool Chains
Gradual tool chains let agents slip into high-privilege territory undetected.

Privilege escalation in LLM agent tool chains does not work the way it works in traditional systems. There is no single exploitable flaw to patch, no misconfigured sudo rule sitting in a file waiting to be found. Instead, privilege builds up gradually, through a string of individually reasonable actions, until an agent ends up somewhere it was never supposed to be. Understanding how that buildup happens, mechanically, is the only real starting point for defending against it.
Classical privilege escalation gives a defender something concrete to find: a SUID binary, an overly permissive IAM policy, a service running as root that should not be. These are discrete, and discrete things show up in a scan. LLM agents break that model in three specific ways. First, the agent picks its own tools at runtime, which means it is effectively choosing its own privilege level on the fly, task by task. Second, untrusted input, a web page, an email, the output of some other tool, arrives through the exact same channel as trusted instructions, with nothing structural separating the two. Third, the agent reasons across multiple steps, chaining together actions that look harmless in isolation but add up to something no single permission check would ever flag. As of 2025, 67% of organizations report having agentic AI running in production, and separate research from the Cloud Security Alliance found that 65% of organizations had an AI agent security incident in just the past twelve months. These paths exist, and they stay invisible until an attacker walks them, which is the actual problem this piece sets out to unpack.
How autonomous tool selection becomes a privilege escalation vector on its own
Give an agent five tools to complete a task and it will sometimes reach for the one with the most access, not because that tool is required, but because it seems like the safer bet for getting the job done. That is over-privileged tool selection, and it is not a design bug sitting in some tool's configuration. It comes out of how the underlying model reasons about finishing a task when it is not fully sure what the task requires.
The ToolPrivBench benchmark, published in June 2026 (arXiv:2606.20023), was built to measure exactly this. Researchers ran it across eight domains and five recurring risk patterns, and the result was blunt: over-privileged tool selection shows up consistently across mainstream LLM agents. Timeouts and partial responses made it worse, pushing agents toward broader-permission tools when the narrow path stalled out. General safety alignment, the kind of training that keeps a model from generating harmful text, did not carry over to least-privilege tool choice. Prompt-level instructions telling the agent to prefer minimal permissions only helped a little.
Every tool negotiation an agent runs, then, is a live escalation event, not just a routine capability decision. If alignment training does not transfer to this problem and prompt instructions only patch it partway, the question becomes what actually can constrain the choice. That question sits unresolved for now, and it sets up the rest of the argument.
What untrusted inputs do to an agent that holds real tool permissions
Prompt injection works because the model has no built-in wall between an instruction from the system and text it happens to be reading. A line buried in a retrieved web page looks, structurally, identical to a line in the system prompt, since the model just sees tokens. Indirect prompt injection takes advantage of this by planting adversarial instructions in emails, web pages, or tool outputs, content the agent runs into during ordinary task execution, never crafted by the actual user sitting at the keyboard.
The numbers on this are not subtle. The HOUYI black-box attack compromised 31 of 36 real-world LLM-integrated applications, an 86.1% success rate, Notion among them. A meta-analysis pulling together 78 studies published between 2021 and 2026 found that attack success rates against current defenses top 85% once the attacker uses adaptive strategies rather than static payloads. Simon Willison named the underlying pattern the "lethal trifecta" in June 2025: private data, untrusted content, and a way to send data out, each fine on its own, together a dependable path to exfiltration.
Scale makes this worse than a single compromised account. AI agents move roughly 16 times more data than a human user doing the same class of work, so one compromised agent is not a one-person incident, it is a high-volume exposure event by default. CrowdStrike's 2026 Global Threat Report, tracking more than 280 adversary groups, documented malicious prompt injection against legitimate AI tools at more than 90 organizations in 2025 alone. This is not a research curiosity anymore; it is operational, happening now.
Injection by itself just redirects what the agent does next. The real damage shows up once that redirection meets a tool with real permissions attached to it, which is where the chain comes in.
How escalation paths form across a chain — the compounding mechanics
Each step an agent takes leaves behind a bit of state that the next step reads and reasons from. Plant a piece of attacker text early in that chain, and it can quietly bias what gets retrieved later, nudge the planning goal sideways, and eventually push a tool call that nothing in the original system prompt ever authorized.
The EchoLeak vulnerability found in Microsoft Copilot (CVE-2025-32711) is a clean example of this pattern. Nothing in the system technically breaks. It moves through a sequence of entirely legitimate states, one after another, until the last one leaks data the earlier ones never touched. STRIDE-style threat modeling misses this because it never asks whether meaning can survive and travel across turns and contexts, which is the actual mechanism at work here.
Laid out as a sequence, it looks like this: input, retrieval bias, planning goal shift, tool invocation, aggregated exfiltration. No single link in that chain is the exploit; the chain itself is the exploit. Traditional access control only ever looks at one request at a time, and it has no way to see the trajectory connecting five requests into a single intent.
Cloud environments show the same shape of problem in lateral movement. A CSA research note on LLM-orchestrated attack chains, published in 2026, describes an agent moving from a compromised workload identity to a fully privileged admin role through a sequence of steps where no individual credential, viewed alone, would ever look like part of a chain.
Some of the more striking evidence comes from research meant to expose exactly this. ChainReactor, presented at USENIX Security 2024 by researchers from King's College London, UCL, UC Santa Barbara, and Vrije Universiteit Amsterdam, used automated AI planning to discover multi-step privilege escalation chains across 504 real Amazon EC2 instances and 177 DigitalOcean instances. It rediscovered known chains and turned up ones nobody had documented before. Separately, hackingBuddyGPT, described by Happe, Kaplan, and Cito in a 2025 Springer paper, had GPT-4-turbo escalate privileges on test systems somewhere between 33% and 83% of the time, through nothing more than iterative reasoning, with no pre-loaded database of known techniques. It figured out the path from first principles, given enough context to work with. A 2026 follow-on by Normann and colleagues pushed this further: a fine-tuned 4-billion-parameter model hit a 95.8% success rate on privilege escalation, nearly matching Claude Opus, at roughly a hundredth of the inference cost. The capability is not staying with the frontier labs; it is moving down.
What ties all of these together is simple: the escalation path never lived inside any one permission, any one tool call, or any one reasoning step. It only existed as the sequence.
Three specific path types that appear repeatedly across real systems
Three patterns recur often enough across real deployments to name individually.
The first is memory poisoning through fragmented reconstruction. Research on this pattern, FragFuse (Rao et al., 2026), shows a prohibited request getting split into pieces across separate interactions, each piece stored in the agent's long-term memory looking completely benign on its own. Later, retrieval stitches the fragments back together into the original disallowed request, which never appeared in full in any single user query. An access control check watching the retrieval query sees nothing wrong, because it is only looking at the query, not the reconstructed meaning sitting behind it. Memory here is not a passive log sitting in the background; it is an active surface where privilege can quietly accumulate over time.
The second is confused deputy behavior, adapted for a world with multiple agents talking to each other. The classic version: an agent invoked by a trusted caller uses the permissions it inherited from that caller to act on behalf of a source nobody actually trusts. In multi-agent systems this becomes collusion, where one agent gets a capability it could never hold on its own by routing the request through a second agent that can. Researchers have demonstrated that a malicious prompt can propagate across a network of connected agents, causing disruption system-wide, one agent's injection becoming every agent's problem. Mandatory access control research addressing this, published on ResearchGate under the title "Taming Various Privilege Escalation in LLM-Based Agent Systems," proposes tying inter-process communications together semantically and checking the whole call-chain at runtime. That kind of check has to live at the middleware or kernel layer, since a prompt cannot enforce it.
The third is tool-chain amplification through protocol-level trust, a risk that emerges when agents can dynamically discover and call tools they were never explicitly wired up to at deployment time, which is genuinely useful and also genuinely risky. Security guidance on agentic protocols flags tool description poisoning as a structural weakness: a malicious tool's own description can talk the agent into calling it, or into calling a legitimate tool in a way that quietly serves the attacker instead of the user. The protocol's design, allowing agents to call tools without human confirmation of each call, is simultaneously the efficiency gain and the exposure. Both tool poisoning and auto-execution abuse have been documented as patterns already showing up in the field, not hypotheticals sitting in a paper.
All three share the same underlying gap: the permission system sees individual, discrete requests, while the agent is actually executing a coherent sequence aimed at a goal. That mismatch is where every one of these paths lives.
Why classical threat modeling frameworks don't see these paths
STRIDE was built to model failures in individual components: a spoofed identity here, a tampered message there. It was never built to track meaning accumulating across a reasoning chain, and it shows. STRIDE has no category for "what happens if future reasoning depends on text an attacker planted three steps ago," which happens to be exactly the question that matters for agentic escalation.
A joint advisory from the Five Eyes intelligence alliance, CISA and the NSA alongside counterparts in Australia, Canada, New Zealand, and the UK, published April 30, 2026, lays out five separate risk categories specific to agentic AI: privilege escalation, design and configuration failures, behavioral misalignment, structural brittleness, and accountability gaps. None of these map cleanly onto STRIDE's categories, because they describe a different kind of system.
MAESTRO, released by the CSA in February 2025, is a better starting point. It splits agent systems into seven layers and pairs with MITRE ATLAS's catalog of 84 documented adversary techniques. Still, it is a cataloging framework, useful for organizing what is known, not a mechanism for catching a novel path as it forms.
The governance numbers make the gap concrete. Netskope's AI Risk and Readiness Report for 2026 found AI tools present in 73% of organizations, while real-time governance enforcement was actually running at only 7% of them. Separately, IBM found that 97% of organizations lacked proper access controls for their AI systems. There is also a timing hole that static threat models simply cannot see into: an agent can write code, call tools, update memory, and change dependencies well before anything reaches a repository checkpoint, and that entire pre-checkpoint window sits outside the scope of any static review process.
This is not a future risk waiting to arrive; it is the current baseline. Catching these paths requires actual runtime visibility into the semantic trajectory an agent is following, not another round of per-request permission checks.
What effective controls for these paths actually require
Least privilege, for an agent, has to mean something different than it means for a human user. It has to be enforced at the moment of tool selection, live, at runtime, not just configured once at deployment and left alone. The reason is straightforward: an agent's actual privilege level is whatever tool it decides to call, not whatever tools happen to be listed as available to it.
Prompt-level instructions do not solve this. ToolPrivBench already showed they only offer limited help against over-privileged selection, and the injection-defense numbers, north of 85% bypass rates under adaptive attack, tell the same story from a different angle. Runtime enforcement needs to check something more specific than "was this agent allowed to call this tool." It needs to check whether the call actually makes sense given the stated task and the sequence of steps that led up to it, the same call-chain verification the mandatory access control research described earlier. Memory retrieval needs to be treated as a genuine trust boundary rather than a passive read, given what FragFuse showed about privilege quietly building up inside stored fragments. Tool descriptions, especially in MCP-style environments, need integrity checks of their own, since an agent's tool-selection reasoning is never more trustworthy than the metadata feeding it.
There is also a verification problem that gets skipped too often. Because these paths only exist as sequences, a fix aimed at one of them has to be checked against the entire set of approved workflows an organization actually runs, not just reviewed in isolation as a diff. Confirming that a safeguard actually stops the attack, without quietly breaking something legitimate, means re-running those workflows end to end. Testing this against a live production agent is a bad idea twice over: it risks disrupting real work, and the results will not reproduce cleanly. What is needed instead is an isolated environment that mirrors the production chain faithfully without being the production chain itself.
None of this counts as proven until it is demonstrated end to end. A claimed escalation path is credible once someone has actually walked it start to finish. A claimed fix is credible once both the attack and the untouched legitimate workflows have been re-run to confirm nothing broke and nothing got through. Because these paths stay invisible until someone walks them, automated tooling that merges a fix silently removes the one moment a human reviewer could actually look at the whole chain. Every safeguard applied to an agentic system should leave behind a record a person can inspect and re-run themselves.
The gaps that remain open even with runtime controls in place
Runtime controls cover the paths that have already been found and verified. They say nothing about the ones nobody has found yet, which, by definition, stay invisible until someone does.
Decommissioning is one of the quieter failure points here. CSA research found that only 21% of organizations have a formal process for retiring agents. An agent that nobody is actively watching anymore, but that still holds live credentials and tool access, is an escalation surface sitting out in the open. Compounding this, the same CSA research found that 82% of organizations have discovered AI agents running in their environment that nobody knew about. Least privilege is meaningless to enforce on an agent whose existence was never logged in the first place.
The capability keeps getting cheaper to run. The Normann 2026 result, a fine-tuned 4-billion-parameter model hitting 95.8% success on privilege escalation at a small fraction of frontier inference cost, means the barrier to executing these chains is dropping fast, for defenders and attackers alike. Runtime controls narrow the window, but they do not close it, and treating them as though they do is its own kind of exposure.


