Indirect Prompt Injection in Agentic Retrieval Pipelines
Attackers exploit agentic systems by embedding malicious instructions in content agents retrieve.

Indirect prompt injection is not a bug someone can patch out of a retrieval-augmented pipeline; it is the predictable cost of building agents that fetch, read, and act on content they did not write. Anyone waiting for a clean fix at a single stage of the pipeline is waiting for something the architecture does not support. The vulnerability sits in the design itself, and it has to be treated that way.
A traditional large language model interaction is narrow: a user types a prompt, the model responds, and the trust boundary sits in one obvious place, whatever the user typed. Agentic systems break that boundary open. A RAG pipeline pulls in documents, web pages, emails, API responses, database rows, and all of it lands in the same context window as the system instructions and the user's own request. The model processes the whole mixture in one reasoning pass. Nothing in a transformer's architecture tags a token as "data to summarize" versus "instruction to obey," so the model reasons over everything in context as if it all carried equal authority, because nothing tells it otherwise.
Persistent state makes this worse over time. Tool outputs and retrieved documents accumulate across turns, so a poisoned artifact pulled in during step one can quietly steer a decision made five steps later, long after the retrieval that caused it has scrolled out of view. The same reach that makes agentic systems worth building, the fact that they can go fetch things and act on what they find, is the structural opening that lets an attacker's text ride in beside legitimate content. What follows walks that pipeline stage by stage: how the injection gets in, how it spreads through context and planning, where it turns into damage nobody can undo, why defenses aimed at one stage keep losing, and what actually has to change.
What indirect prompt injection is and how it differs from direct prompt attacks
Direct prompt injection is the version most people picture: an attacker types something adversarial straight into the chat box, trying to override the system prompt. It requires access to the interface, and it shows up in the conversation log where anyone can find it.
Indirect prompt injection does not need the interface at all. The malicious instruction sits inside content the agent is going to retrieve on its own: a webpage, a shared document, an email, a code comment. It reaches the model through a tool call or a retrieval step, never through anything the user or operator typed. Under any static inspection, that content reads like ordinary text; it only does anything once the model processes it during a live run, which is exactly what makes scanning documents ahead of time useless against it.
When it lands, the agent quietly deviates from the task it was given: it picks a different tool, takes an action nobody asked for, sends data somewhere it should not, while the conversation on the user's screen looks completely ordinary.
This is where agentic systems part ways with stateless chatbots. An agent acts rather than just answering, so a hijacked agent can call APIs, write files, send messages. An agent retrieves content on its own initiative, so the attacker's payload enters context with no human reviewing it first. And an agent plans across multiple steps, so one injected line can redirect an entire workflow instead of a single reply. OWASP's Top 10 for LLM Applications, in its 2025 edition, ranks prompt injection as the number one risk and frames it specifically as an exploitation of the model's inability to tell trusted instructions apart from untrusted data. That is a structural fact about how these systems process context, not a misconfiguration waiting to be patched. Jailbreaks argue the model out of its system prompt directly; indirect injection does not argue with anything. It simply hands the model content the model was already built to read and trust.
The three conditions that make an agent fully exploitable
Simon Willison named this the "lethal trifecta" in 2025, and the label earns its keep because it turns a vague sense of danger into a checklist. Three conditions have to sit together before indirect injection stops being a curiosity and becomes a real breach.
First, the agent needs access to private data, files, messages, credentials, internal records, something worth stealing that the attacker cannot reach directly. Second, it needs exposure to untrusted content it will actually fetch or process: web pages, public documents, emails, comments on a shared repository. Third, it needs a way to communicate externally, an HTTP call, an outbound email, an API write, anything that lets a result leave the system and reach the attacker or otherwise change something in the world.
Remove any one of the three and the attack path collapses. No private data, nothing worth exfiltrating. No untrusted content, no way to inject an instruction in the first place. No outbound channel, and even a fully hijacked agent has nowhere to send what it stole. The uncomfortable part is how ordinary this combination has become: most enterprise RAG assistants and tool-using agents satisfy all three conditions simply by doing their job, since a useful agent generally needs a private knowledge base, the ability to pull in outside content, and tool calls that write or send something. Treating the trifecta as a rare edge case is the mistake; it is closer to the default configuration of any agent worth deploying. The diagnostic question worth carrying into the rest of the pipeline is which stage introduces each condition, and whether it needs to.
How the injection enters: retrieval and tool-call ingestion as the first chokepoint
Every agentic pipeline has ingestion points: web search results, document loaders, email readers, API responses, outputs from Model Context Protocol services, retrievals from a vector database. Each one is a place where content outside the operator's control crosses into the model's reasoning.
The RAG case shows the pattern cleanly. An attacker seeds a knowledge base with text that reads as coherent, on-topic content but is actually written to surface for specific queries and to carry an instruction once it does. Research on this attack class has formally demonstrated exactly this: crafted text engineered to get retrieved on target queries and redirect the model's output once it does. To any indexing pipeline running upstream, the poisoned chunk looks like an ordinary document chunk. Nothing about it flags as malicious.
The live-fetch case asks even less of the attacker. Plant instructions in a public web page, a shared document, an email, anything the agent will go fetch on its own, and there is no need for write access to the target's systems at all. Publishing content somewhere the agent will retrieve it is enough, and because the payload is natural language rather than executable code, antivirus tools, firewalls, and static scanners have nothing to flag.
Model Context Protocol, a standard for connecting agents to tools and services, widened this surface considerably as adoption spread. Agents built on MCP ingest not just retrieved documents but tool descriptions and capability metadata as working context, and each connected service becomes a potential injection point: the tool description itself can carry an adversarial instruction that shapes planning before any retrieval even happens. The injection surface extends across the ordinary web, not confined to sites anyone would have flagged as suspicious beforehand.
How injected content propagates through context and shapes planning
Once retrieved content lands in the context window, it sits next to the system instructions with no internal wall separating the two. Whatever tool the model picks next, whatever query it issues, however it reads an intermediate result, gets shaped by everything currently in context, injected material included.
In pipelines that carry state across turns, this compounds. A poisoned chunk pulled in during step one can steer what the agent decides in step three, or step five, or step seven, well after the retrieval that caused it has dropped out of view. The injected instruction does not need to argue with the system prompt or override it outright; it only needs to make some other action look like the correct next step given whatever goal the agent is currently pursuing.
The newest wrinkle sits in the memory layer. Research in the 2025 to 2026 window showed adversarial instructions embedded in memory entries surviving across sessions entirely, redirecting tool selection in later interactions triggered by prompts that have nothing to do with the original injection. That stretches the attack horizon from the current conversation to an indefinite number of future ones; the injected instruction becomes a standing directive sitting in memory, waiting for its moment. It also breaks a comforting assumption a lot of system designers have leaned on, that session isolation caps the damage. It does not.
Underneath all of this sits one root cause: the agent was built to treat retrieved content as legitimate input for its own planning, and indirect injection works by making an attacker's instruction indistinguishable from a fact the agent just looked up honestly. The danger was never really about any single poisoned document. It is about an architecture that grants retrieved content a default trust it never earned.
Where the damage happens: tool execution and external action as the terminal stage
Tool-augmented agents do not stop at producing text. They call APIs, read and write files, send messages, run code, reach into other services. Once planning has been hijacked upstream, every tool call downstream executes under the attacker's redirected intent, and the user watching the session sees nothing that looks abnormal.
A few patterns recur at this terminal stage. Data exfiltration: the agent reads something private and transmits it out through an API call, an email, or a request to an attacker's endpoint. Credential harvesting: tokens, keys, session credentials the agent happens to have visibility into get pulled and forwarded. Lateral action: the agent gets pushed into invoking other tools that extend the attacker's reach past the original session. Content manipulation: the agent writes attacker-controlled text into a shared system, a wiki, a pull request, a message thread, that other people or other agents will later read and trust without a second thought.
Documented production incidents illustrate this chain running end to end. In cases that have reached public disclosure, crafted content delivered through ordinary channels caused AI assistants to access private organizational data and route it externally, with every step written in plain natural language. Antivirus tools, firewalls, static scanners: none of them had anything to catch, because there was no code to catch.
Other incidents follow the same shape at smaller scale. Across tool-augmented assistants, injected instructions have caused agents to surface private content, forward sensitive credentials, push attacker-controlled text into shared systems, and take actions with real financial or operational consequences. This is where an architectural risk that reads as abstract on paper turns into a specific, named, dated organizational loss.
Why defenses applied at a single pipeline stage consistently fail
The obvious first instinct is to filter content before it reaches the model, catch the bad text on the way in. That instinct hits a hard limit fast: the injection is written in ordinary language, not a code pattern with a signature to match against.
Classifier-based detection has been tried at real scale, and it lost. Documented production incidents have shown injections bypassing enterprise-grade classifier defenses, about as strong a real-world test as such a defense will ever face. Research testing published defenses under adaptive attack conditions — where the attacker can study and adjust to the defense — has consistently found that attacker success remains high. The defender always moves first and publishes; the attacker studies what got published and works around it.
Prompt-level hardening, telling the model in its own instructions to ignore anything that looks like an embedded command, works fine against clumsy attempts and falls apart against anything crafted to present itself as a natural continuation of the agent's own goal. The underlying problem was never the wording of the defense. It is that the model has no reliable way to tell "this is data I am summarizing" apart from "this is an instruction I am supposed to follow." That is the exact architectural gap indirect injection lives inside.
Session isolation, the assumption that whatever happens in one conversation stays contained to that conversation, does not hold either. Research on memory-layer attacks showed injected instructions riding across session boundaries inside the memory subsystem itself, with production exploits carrying CVSS scores above 9.0. Isolation shrinks the blast radius. It does not close the path.
Each of these defenses fails for the same underlying reason: it addresses one stage and misses the others. Filtering at ingestion misses whatever happens once content propagates through context. Guarding the tool call misses the planning hijack that already happened upstream of it. Hardening the model layer misses memory persistence that survives past the session entirely. Benchmark evaluations of published defenses still show even the best-performing approaches leaving a substantial share of injection scenarios unblocked, and separate red-team testing from NIST on novel attacks found task-hijack rates well above what earlier baseline estimates had predicted. None of this means defenses are worthless; it means no single control, applied at a single stage, is enough on its own. Any defense worth deploying has to be tested against the real, adaptive attack, not a tame version drawn up in advance and never pressure-tested again.
What a pipeline-aware defense posture actually requires
Because the attack moves through several stages rather than living in one, a serious defense has to cover every point where trust changes hands, not just the point where outside content first arrives. Picking one stage and hardening it well is not a defense strategy; it is a false sense of one.
At ingestion, that means treating every piece of externally retrieved content as untrusted by default, structurally, in how the system is built, not as a policy written down somewhere and hoped for. Where the architecture allows it, retrieved content should stay visibly separate from the instructional context sitting beside it, rather than blending into one undifferentiated block of text.
At the planning stage, the sharpest lever is permission scope. An agent that cannot read private data or reach an external channel cannot complete the lethal trifecta, no matter how thoroughly its reasoning gets hijacked upstream. Every tool call should be scoped to the narrowest set of actions the approved workflow actually needs; every extra capability sitting unused is one more leg of the trifecta left open for no reason at all.
At the point of execution, tool calls need runtime checks against what the approved workflow is actually supposed to look like, flagging a deviation before it fires rather than logging it afterward. High-consequence actions, touching credentials, writing outside the system, moving money, need a human checkpoint in the loop, no exceptions. In each of these cases, the actions involved were ones that a human reviewer, given the chance to see them before execution, would have stopped cold. That is the standard worth building toward, not a filter that promises to catch the bad text before it arrives.


