Threat Hunting in Environments With Deployed AI Agents
AI agents expand their attack surface in real time, outpacing traditional threat-hunting playbooks.

Threat hunting inside environments running deployed AI agents requires a different set of instincts than the ones built for static infrastructure, because the agents themselves keep redrawing the map of what counts as an attack surface. AI tools already sit inside 73% of organizations, yet real-time governance enforcement, the kind that would actually catch misuse as it happens, has reached only 7% of them https://www.augmentcode.com/guides/ai-agentic-threat-modeling. That gap between adoption and control is the whole story in a single measurable distance. It's the whole problem in two numbers.
IBM's Cost of a Data Breach Report found that 13% of organizations had already reported a breach touching an AI model or application, and 97% of organizations lacked proper AI access controls https://www.augmentcode.com/guides/ai-agentic-threat-modeling. Those two figures sitting next to each other say something plain: the exposure is not theoretical, and neither is the absence of the guardrails that would normally catch it. Meanwhile, the people closest to this problem expect it to get worse before it gets better. Nearly half of surveyed practitioners, 48%, believe agentic AI will be the top attack vector for cybercriminals and nation-state actors by the end of 2026 https://www.augmentcode.com/guides/ai-agentic-threat-modeling https://www.darkreading.com/threat-intelligence/2026-agentic-ai-attack-surface-poster-child.
So the situation facing threat hunters right now is not abstract risk modeling. Agents are live in production, the controls around them are thin, and the people whose job is to find the intrusion before it becomes a headline are being handed a surface with no established playbook. Agents don't hold still long enough for that assumption to work.
Runtime attack surface expansion in agentic systems
A conventional application has a fixed attack surface the moment it ships. Whatever ports are open, whatever APIs are exposed, whatever code paths exist forms the surface, and it stays roughly constant until the next release. An autonomous agent doesn't work that way. Its surface grows continuously, every time it negotiates a new tool, writes to a memory store, or exchanges a message with another agent, none of which were fully specified when the system was designed.
Three mechanisms drive that expansion, and hunters need to understand each one on its own terms. Tool calls are the most visible: agents invoke external APIs, execute code, touch file systems, reach into internal services, and every single one of those invocations creates a new trust boundary at the moment it happens, not back when the system was architected. Memory writes are quieter but arguably more dangerous, because what an agent writes to its memory today shapes how it reasons and acts in a future session, and a threat can be planted with no immediate observable event to flag it. And in multi-agent pipelines, one agent's output becomes another agent's instruction, so a single compromised upstream agent can push malicious intent downstream through inter-agent messages without the original attacker doing anything further after the initial injection.
This is exactly where older threat modeling frameworks start to fail, and the failures are documented. Take goal hijacking: an agent uses its own legitimate capabilities against the operator who deployed it. STRIDE has no category for that, because STRIDE was built to describe spoofing, tampering, and denial of service against a system convinced to work against its own owner using tools it was authorized to use all along. The OWASP Agentic Top 10 had to create a new entry for it, ASI01, precisely because nothing in the older taxonomy covered it.
Component-level security review runs into the same wall. EchoLeak, the zero-click prompt injection vulnerability found in Microsoft 365 Copilot, exfiltrated internal files with no user interaction at all, and every individual component involved could pass a security review in isolation while the attack path itself remained wide open. Component-by-component analysis checks each piece for correct behavior on its own, but the exploit lives in how the pieces connect, not in any single piece. Trust boundaries inside an agentic system are dynamic rather than fixed at a perimeter; an instruction like "never share sensitive data" functions as a probabilistic nudge on model behavior, the way a firewall rule is enforced as an actual control. And persistence through memory has no classical category. The ATFAA framework had to invent one, defining temporal persistence threats as a domain STRIDE simply never addressed.
Security researchers now track the evolution of these attacks across escalating levels, and the progression is instructive. Direct prompt injection is at the bottom, where an attacker feeds the model instructions meant to hijack its stated goal or leak its system prompt. Jailbreaking through crafted adversarial input follows. Above that sits indirect prompt injection via third-party content, the level responsible for incidents like PoisonedRAG and the Slack AI exfiltration case, where the malicious instruction never comes from the user at all, it comes from something the agent reads. Higher still is agentic exploitation through tool chains and MCP servers. Tool poisoning, EchoLeak, and the Cursor remote code execution case all live there. Multi-agent system attacks sit at the emerging edge, propagating through inter-agent messages with the potential to cascade across organizational boundaries.
The telemetry layers hunters must learn to read in an agentic environment
None of this means network logs, endpoint telemetry, and SIEM alerts stop mattering. They remain necessary for watching the infrastructure layer, but they were never built to see inside what's actually happening in an agent's reasoning, memory, or tool execution, and that blind spot is where the interesting attacks now live.
Picture the agent as running on a stack with its own layers, each demanding its own telemetry. The cognitive layer covers model reasoning traces, prompt and completion logs, and system prompt version history, and goal hijacking and indirect injection first appear there if anyone is looking. The memory layer covers vector store write logs, episodic memory mutation events, and embedding change records, and this is specifically where MemoryGraft-style attacks embed themselves across sessions without tripping anything visible at the network layer. The tool execution layer covers MCP server call logs (full argument and result payloads, not just whether a call succeeded), tool authorization records, and credential invocation audit trails, and this is the layer where exfiltration of actual data physically occurs. Above all of that sits the inter-agent message layer, made up of orchestrator routing logs, the instruction payloads agents send each other, and shared memory read and write events.
A lot of deployed guardrails quietly fail here. Runtime guardrails that only scan the incoming user prompt never see what comes back from a tool call or a database query, so if an MCP server returns sensitive data, input filtering has nothing to catch, because the filter was watching the wrong direction of traffic. Hunters need visibility on data moving between agents and servers as well as on what a user typed.
There's also a structural gap. A plain chronological log can tell a hunter that a document was retrieved, a tool was called, and an answer got produced, in that order. It cannot tell them which document influenced which tool argument, or which tool result changed which downstream reasoning step. That requires a provenance graph, not an event sequence, because runtime guardrails that only scan the user prompt miss what returns from tools and databases (if an MCP server returns sensitive data, input filtering does not catch it, so hunters need telemetry on data moving between agents and servers, not just user-facing prompts).
Following a multi-step attack chain: prompt injection as the entry point hunters most often miss
Prompt injection tops the OWASP Top 10 for LLM Applications 2025 as LLM01, with attack success rates in agentic systems reaching 84% and production exploits tied to it carrying CVSS scores above 9.0 https://www.augmentcode.com/guides/ai-agentic-threat-modeling. Those aren't lab numbers describing a hypothetical risk. They describe a technique already working against deployed systems at a rate that should worry anyone running an agent with real permissions.
Security researchers describe the conditions for a successful attack using what's come to be called the lethal trifecta: an agent needs access to private data, exposure to untrusted content, and the ability to communicate externally, all three at once, before the attack path closes. Removing any one leg breaks the chain. That gives hunters a genuinely useful triage filter. Rank agents by how many legs of the trifecta they satisfy, and the highest-priority hunting targets sort themselves.
Indirect injection is the variant most hunting programs still can't see, mostly because the instruction never touches the interface anyone is monitoring. In January 2025, researchers demonstrated exactly this against a major enterprise RAG system: malicious instructions embedded inside a publicly accessible document caused the AI to exfiltrate data and rewrite its own system prompt, with zero user action required once the document had been planted. The user never typed anything malicious. They just asked a normal question, and the poisoned document did the rest.
That threat is scaling. Google researchers monitoring the open web found a 32% increase in malicious prompt injection payloads embedded in web content between November 2025 and February 2026 https://www.augmentcode.com/guides/ai-agentic-threat-modeling. The content an agent reads to do its job is actively weaponized terrain. It's actively weaponized terrain, and treating it as trustworthy background material is no longer defensible.
Goal hijacking sits above indirect injection as the more damaging variant, because it doesn't just extract one document's worth of data, it redirects an agent's entire objective across a multi-step pipeline. One compromised agent poisons shared memory, then manipulates the decisions an orchestrator hands to downstream agents, and the original injection point may never register at the network layer. By the time anyone notices the output looks wrong, the actual entry point could be several hops and several sessions back.
Threats that accumulate across sessions: memory poisoning and temporal persistence
Traditional threat hunting looks for a signal that fires, such as a spike, an anomaly, or a moment something crosses a threshold. Memory poisoning doesn't work that way. It plants a condition that quietly shapes future behavior, with no single observable event marking the moment of attack. A hunter looking for a firing signal will simply never find one.
Research on MemoryGraft describes exactly this category: malicious content corrupts an agent's persistent memory store, and the corrupted "experience" that gets implanted looks, at the storage layer, indistinguishable from anything the agent legitimately learned. There's no flag, no anomaly score, no obvious tell. It's memory that happens to be lying.
Retrieval-augmented generation systems face the same problem from a different angle. Targeted RAG poisoning can steer an agent's answers using a surprisingly small number of planted documents, and the academic literature synthesizing this evidence, including the SoK work from Dehghantanha and Homayoun, confirms that meaningful steering doesn't require flooding a corpus, just a handful of well-placed documents. That makes the retrieval layer a double-edged asset: it's an intelligence source the agent depends on, and simultaneously a persistent attack surface an adversary only needs to touch once.
Tool poisoning research reinforces the same pattern at the execution layer. Studies on this class of attack report success rates against AI agents ranging from 36.5% up to over 80% depending on conditions, with agent refusal rates sometimes dropping under 3%. Hunters relying purely on inspecting an agent's final output will not catch this. The manipulation happens upstream of the output, in the tool description or the tool's returned payload, long before anything reaches a screen someone is watching.
Adapting hunting hypotheses and workflows to the agentic threat model
Hypothesis-driven hunting doesn't need to be thrown out. It needs to be re-pointed. Instead of building hypotheses around IOC matches or single-event anomalies, hunters need hypotheses structured around attack paths that cross the cognitive, memory, and tool layers of the stack, because that's where these attacks actually live.
MAESTRO's seven-layer framework gives that effort a usable scaffold. At Layer 1, Foundation Models, hunters look for prompt injection signatures inside completion logs and for patterns where the model appears to override its own instructions unexpectedly. At Layer 2, Data Operations, the hunt turns to anomalous vector store writes, embedding drift, and signs of log tampering inside retrieval pipelines. At Layer 3, Agent Frameworks, hunters watch for tool misuse, unexplained goal shifts inside reasoning traces, and plugin execution happening outside approved workflows. Layer 4, Deployment and Infrastructure, covers infrastructure tool exploitation and resource exhaustion patterns consistent with an agent being abused rather than simply busy. Layer 5, Evaluation and Observability, is where hunters look for log evasion and attribution gaps that appear specifically during multi-agent handoffs, the exact seams where accountability tends to disappear.
Layer 7, Agent Tools and Integrations, deserves particular attention. None of those map cleanly onto instincts built from a decade of endpoint and network security work.
The most important lesson from MAESTRO's follow-on research is that the most dangerous attack paths rarely stay inside one layer. They start at Layer 1, with a prompt injection, and cascade through Layer 4 and beyond, so a hunter chasing a single thread has to be willing to follow it across layer boundaries rather than stopping once the trail leaves their usual area of ownership.
MITRE ATLAS offers a concrete starting point for building out a hypothesis library rather than starting from a blank page. The Poisoned Postmark MCP email exfiltration case, catalogued as AML.CS0053, shows how a poisoned tool description can turn an email integration into an exfiltration channel. The Bing Chat indirect prompt injection case, AML.CS0020, remains one of the clearest public examples of third-party content hijacking a model's behavior. And the LAMEHUG malware, attributed to the Russian state-backed group APT28 and catalogued as AML.CS0044, demonstrates that this isn't confined to opportunistic criminal activity. Nation-state actors are already building tooling around these same weaknesses.
None of these cases are hypothetical scenarios dreamed up to stress-test a framework. They're documented incidents, and each one maps to a specific layer, a specific mechanism, and a specific gap in the telemetry that traditional hunting workflows were never built to fill. Building a hunting program around agentic deployments starts with accepting that the surface itself moves, and that the tools built for a surface that stood still will always be one step behind it. MITRE ATLAS documented 84 adversary techniques against AI systems (MITRE ATLAS, https://www.augmentcode.com/guides/ai-agentic-threat-modeling).


