Est.

Provenance Tracking for AI Agent Tool Call Sequences

Tracking where agent actions originate stops harmful tool chains before they cascade.

Senior Writer · · 10 min read
Cover illustration for “Provenance Tracking for AI Agent Tool Call Sequences”
Runtime Controls · September 24, 2026 · 10 min read · 2,224 words

AI agents no longer just answer questions. They act: they call tools, write files, send messages, and hand results forward to the next step in a chain they generated themselves. Provenance tracking is the discipline of recording exactly where each of those steps came from and what fed into it, so that a harmful outcome can be traced back to its cause and, ideally, stopped before it completes. This piece explains what that means in practice, why the multi-step nature of agent execution makes it unavoidable, and how the current generation of research systems tries to build it.

How agentic execution expands the attack surface across tool calls in a sequence

A stateless API call does one thing and returns. An agent tool call does that, then changes the terrain for whatever comes next: it might unlock a credential, write to memory, or trigger a call to a separate service that the agent didn't have access to a moment earlier. That means the attack surface isn't something you can diagram at design time and call finished. It grows at runtime, one tool call at a time, shaped by whatever the agent reads, writes, and passes along.

The Cloud Security Alliance's MAESTRO framework breaks agentic risk into seven layers, and Layer 7, the agent ecosystem layer, is where conventional security thinking stops applying cleanly. Risk at that layer doesn't come from any single tool or model. It comes from composition: how tools, memory, and outside integrations combine across a sequence that no one designed end to end.

A few attack patterns exploit this directly. Indirect prompt injection is the clearest case: the attacker never touches the agent's input at all, but instead poisons a document, webpage, or tool output that the agent later retrieves mid-sequence, and the injected instruction rides in disguised as data. Tool poisoning works a level deeper, hiding instructions inside a tool's own schema so they fire the moment the agent calls that tool; research found that this kind of poisoning slips past output-based safety checks in 95% of tested cases. Once one step is compromised, the effect doesn't stay contained. It cascades forward through every downstream call that trusts the tainted output.

EchoLeak, the zero-click vulnerability disclosed in Microsoft 365 Copilot, is the sharpest illustration of what this looks like in production. An indirect prompt injection exfiltrated internal files with no user interaction at all, the entire chain completing before anyone was in a position to notice. Harm doesn't happen at the instruction; it happens several tool calls downstream of it, and by the time it's visible, it's already done.

Provenance tracking in agent tool call sequences

Provenance tracking, applied to agent tool calls, records the full path an action took: which tool got selected, which arguments were passed to it, what came back, and how that result shaped the reasoning or actions that followed. It's a different question than the one ordinary monitoring asks. Monitoring asks whether a call succeeded. Provenance asks whether the call was justified in the first place, its arguments trace back to something trustworthy, and its output is being used the way it should be.

Inside a single tool call, four objects matter for provenance purposes: the tool that got selected and its schema, the specific argument values and where they originated, the result the call produced, and whoever consumes that result downstream. Of these four, argument origin deserves the most attention, because a completely legitimate tool becomes dangerous the moment its parameters come from an untrusted or corrupted source. The tool was never the problem. The argument was.

How provenance graphs represent the structure of a multi-step agent execution

Formally, provenance in an agentic workflow gets modeled as a directed, attributed graph, one that encodes both the order things happened in and the meaning of the relationships between them. Graphectory, one system built for this purpose, defines a typed structure with three kinds of nodes: prompts, actions (the actual tool or API calls), and validation steps, each tagged with an identifier and metadata like model version, tool name, file path, and the arguments passed. Edges in that graph split into two kinds too: temporal edges that preserve chronological order, and semantic edges that capture derivation, meaning which value came from which source, or which action used which resource.

The W3C's PROV-DM standard supplies a general vocabulary for this kind of thing, built around entities, activities, agents, and the derivation relations between them. PROV-AGENT adapts that vocabulary specifically for agentic workflows, modeling prompts, model responses, decisions, tool interactions, and the surrounding workflow context as a connected structure rather than a flat sequence.

What a graph like this shows that a flat log never will: which prompt or retrieved document a given argument value actually traces back to, which tool call consumed that value, which later calls depended on that call's output, and, critically, where in the sequence a trust boundary got crossed. A log tells you a call happened. A graph tells you why it was allowed to.

How provenance is instrumented at runtime in practice

Three mechanisms do most of the actual capturing. Decorators wrap individual agent tool functions, PROV-AGENT's @flowcept_agent_tool being one named example, intercepting each call to record its arguments and return value, attach metadata, and emit a provenance record. Wrappers sit around the LLM API call itself, recording the prompt and response as provenance nodes at the moment of invocation. Observability hooks tie into distributed compute frameworks, letting provenance signals get ingested in real time across multi-agent or distributed setups where a single wrapper wouldn't see the whole picture.

What all three share is a commitment to dynamic, runtime capture rather than a static map of what an agent could theoretically do. The record reflects what actually happened, with the actual values involved.

Systems in this space don't all optimize for the same thing. AgentOps and AgentTrace lean toward runtime trace capture and general observability. Agent-Sentry narrows in specifically on whether sensitive tool arguments were influenced by untrusted sources. FIDES formalizes an information-flow control model that enforces trust distinctions as execution proceeds. NeuroTaint traces taint influence from an input all the way to the action it eventually shapes.

At what point in the pipeline does the provenance record actually get written, and is it written before the call executes or after? That distinction is not a technicality. A record written after the fact supports audit, nothing more. Only a record written before the call, and checked against a policy, can stop the call from happening.

Why invocation-level controls fail and argument-level enforcement is necessary

Research from 2026 converges on a specific diagnosis: most existing defenses operate at the wrong level of granularity. They decide whether an agent is allowed to call a tool at all, but they never inspect whether the particular argument values going into that call deserve trust. That's a meaningful gap, and it has a concrete failure mode. An agent retrieves a document, a perfectly legitimate action, then takes a value out of that document and passes it as an argument to a write or send operation. The retrieval was fine. The argument wasn't. An invocation-level control sees only "the agent called a write tool" and has no way to tell those two situations apart.

Fan et al. tested this gap directly with a system called PACT (Provenance-Aware Call-time Trust), pitting it against competing defenses on mixed-trust diagnostic scenarios. Vanilla execution, no defense at all, preserved full utility while providing no security, which is what you'd expect from doing nothing. FIDES achieved partial results on both dimensions without excelling at either. CaMeL flipped that tradeoff hard, reaching 100% security but at substantially reduced utility. PACT's L2 configuration was the outlier: 100% utility and 100% security, with zero false positives and zero false negatives on that benchmark.

The pattern held up at larger scale, too. Across full AgentDojo deployments spanning five different models, PACT reached 100% security on the three strongest models while still recovering between 38.1% and 46.4% of utility, meaningfully ahead of CaMeL at matched security levels. Enforcement that can't see individual arguments has to choose between blocking too much and blocking too little.

Competing architectural approaches to provenance-aware enforcement, and what each trades away

Work published between 2024 and 2026 keeps landing on the same core strategy: put security enforcement outside the model, as a deterministic policy layer that mediates what the agent is allowed to do, rather than hoping the model itself refuses a malicious instruction when it counts. Models are trained to be helpful. Policy layers don't have that conflict of interest.

Four architectural families have emerged from this, and each gives something up. Information-flow control systems, CaMeL and FIDES among them, separate trusted control flow from untrusted data and enforce labels at execution time; this works, but it can over-restrict agents that have a legitimate reason to act on retrieved content, and this restriction is visible as lost utility in the PACT comparison above. Systems like NeuroTaint and Agent-Sentry propagate influence labels from source to sink across the model's reasoning and the surrounding code, finer-grained than a simple invocation check, but the labeling and tracking work adds real overhead. Argument-level provenance enforcement, PACT's approach, evaluates the specific values going into a call at the moment of the call, which is what let it hit both high utility and high security on the tested benchmark, though it demands instrumentation infrastructure most teams don't have sitting around already. Runtime reference monitors such as RTBAS and FORGE interpose on agent actions as a single enforcement point; how well they work depends entirely on how expressive their policies are and how deeply they can actually inspect where an argument came from.

Roughly 70% of the mitigations catalogued in MITRE's ATLAS framework map onto security controls that already exist in some form. Provenance-aware enforcement is one of the pieces that doesn't, which is a fair measure of how new this problem actually is. No architecture on the table today handles every threat class at production scale without giving something up. The tradeoffs between how fine-grained the enforcement is, how much overhead it adds, and how much it actually covers are design decisions a team has to make deliberately, not details that resolve themselves during implementation.

The granularity vs. overhead tradeoff in production deployments

Fine-grained, argument-level provenance is what makes verification, debugging, recovery, and audit actually possible after something goes wrong. It also costs more: more storage, more logging complexity, more privacy exposure, more annotation work up front. Coarse-grained provenance is cheaper to collect, but when something breaks, it often can't tell you which memory item caused the error, which piece of retrieved evidence supported a bad claim, or which specific argument crossed a trust boundary it shouldn't have.

The tradeoff plays out across four dimensions in practice. Storage is the most obvious: capturing every argument across every call in a long agent sequence adds up fast, and retention policy has to weigh audit value against the cost of keeping it all. Privacy cuts the other way, since argument values often contain the exact things you don't want sitting in a log verbatim, credentials, personal data, internal content, and a provenance record that captures those faithfully becomes its own liability. Annotation burden is the one people underestimate: several approaches require developers to label the trust level of data sources and tool schemas by hand, and that's not a setup task you do once, it's ongoing engineering overhead. Latency rounds it out, since checking a policy before every tool call adds a step before execution; the delay is bounded, but it's never zero.

None of that argues for uniform treatment across every tool an agent has access to. High-privilege tools, the ones with write access, credential use, or the ability to reach outside the system, justify argument-level capture even at real cost. Read-only tools with a narrow, bounded scope can usually get by with something coarser. The sharpest way to draw that line is the "lethal trifecta": an agent that has access to private data, is exposed to untrusted content, and can communicate externally. Any agent sitting at that intersection is exploitable by construction, and that's the profile that demands the finest-grained provenance available.

What a provenance record must contain to support audit and enforcement

A record built only for after-the-fact review and a record built to stop a bad call in progress aren't the same artifact, though they overlap heavily. At minimum, a provenance record needs the identity of the tool called and its schema, the argument values passed and a trace back to where each one originated, the result the call returned, and a link forward to whatever downstream step consumed that result. Without the origin trace, you're back to a flat log: it tells you what happened, not whether it should have.

The distinction that actually separates audit from enforcement is timing. A record generated after a call executes can tell an investigator, days later, exactly which document supplied the argument that triggered a bad outcome. A record generated and checked before the call executes is the only kind that can refuse the call. Both have a place. But conflating them, treating a good audit trail as though it were a security control, is the mistake that leaves the gap PACT and its contemporaries were built to close.

Diagram: PACT vs. Competing Defenses: Security and Utility Scores. Visualizes: Show how four defense configurations trade off security against utility on the benchmark from Fan et al.'s PACT study.

Sources

  1. AI/Agentic Threat Modeling: Securing Systems That Include Agents
  2. impact.ornl.gov
  3. arxiv.org
Filed underRuntime Controls

More in Runtime Controls