An IETF Internet-Draft of 25 September 2026 proposes a JSON-based, hash-chained audit trail format for autonomous AI agents, with recording before an action and an independent recorder, but it remains a working document and not an adopted standard. Reliable reconstruction requires you to record execution facts, evidence relations and reasoning provenance separately.
What exactly does the new IETF Internet-Draft propose for audit trails for autonomous AI agents?
The working draft draft-sharif-agent-audit-trail-05 from the Internet Engineering Task Force describes a fixed format for recording which agent, under which version and with which authorisation, carried out an action. The document is expressly an Internet-Draft and must be read as work in progress: it is not an established standard and not proof of compliance.
The proposal goes further than ordinary telemetry. The main components:
- Agent identity and sessions: which agent, which version, which session and which trust level.
- Recording before the action: a record that comes into being before an agent acts, not only afterwards.
- Hash chaining: records are cryptographically chained together, so that changes made afterwards become visible.
- An independent recorder: a component outside the agent process that records the events.
- Reconstruction after interruption: so-called prior-generation tails to stitch a chain back together after a failure.
Why is logging beforehand different from logging afterwards, and why can one type of log not reconstruct everything?
Recording afterwards what an agent did does not prove that an authorisation or policy check existed before that action. That is precisely why the distinction between pre-execution and post-execution records in the draft standard matters: a record that comes into being before the action shows that the check took place while it could still intervene, not that it was written in retroactively.
The academic literature points to a second limit. In our reading of the sources there are three reconstruction levels that complement each other and do not follow from one type of log:
- Execution facts: agent version, tool call, outcome and session order. The survey From Agent Traces to Trust by researchers on arXiv states that the final answer alone does not explain how an agent arrived at an outcome.
- Evidence relations: which sources, which memory and which tool responses supported an action. The same survey describes these as a typed graph of relations between action and evidence.
- Reasoning provenance: intent, observation, inference, plan change and delegation. The paper Reasoning Provenance for Autonomous AI Agents states explicitly that this provenance cannot be reliably reconstructed from existing execution traces.
The consequence that none of the sources spells out aloud we name here as an editorial conclusion: whoever keeps only existing traces can technically replay a session without ever being able to demonstrate why the agent chose a particular action. That gap must be recorded separately, or it does not exist.
What limits does this evidence have, and what remains unprovable?
Hash chaining detects whether records have been changed afterwards, but it does not prove that the recorder saw all the events. An incomplete chain that is internally consistent looks reliable and is not. The paper NovaFabric on tamper-evident, replayable execution capsules is honest on this point: the evidence capture is not complete, simulated replay does not automatically replace external tool responses, and the guarantees apply only within an explicitly trusted computing base.
In our estimation, a sober lesson follows from that. Three things remain unprovable as long as you do not organise them yourself:
- Whether the recorder was complete, rather than merely internally consistent.
- Whether a replayed session reproduces the actual external effects, or only the model steps.
- Whether a reconstructable decision was also a correct decision. Reconstruction is no justification.
What does this mean for directors, lawyers and CISOs who work with sensitive information?
Our analysis: because an independent recorder outside the agent process is central in the draft standard, a log file that the agent keeps about itself can lose its evidential value as soon as that agent itself fails or is manipulated; that is why we advise that a CISO assess, per autonomous workflow, whether the recording should be physically and organisationally separated from the system it records. Because recording before the action demonstrates that an authorisation or policy check existed while it could still intervene, a log written only afterwards is weaker evidence for a lawyer; that is why it is advisable to include in supplier contracts a pre-action authorisation record with agent version and model identity as a delivery requirement. Our analysis: because reasoning provenance cannot, according to the research, be reliably derived from ordinary traces, a director may afterwards be unable to explain convincingly why an agent acted if intent, plan change and delegation chain are not recorded separately; that is why the officer responsible for a high-risk workflow must now decide which of the three levels to keep, instead of assuming that one system covers everything. Because hash chaining demonstrates change but not completeness, the risk remains that a crucial event was simply never recorded; that is why a periodic integrity and completeness check of the chain belongs among the standing controls, not among a one-off acceptance test.
As a summary of that analysis: a workable minimum to be able to account for an autonomous AI action later includes an independent recorder, version and model identity, an authorisation record before the action, the data and tool access used, the evidence relations, the human escalations, the interruptions and a check on the integrity of the chain. This connects to earlier work on recording mandate, responsible officer and remediation per action, on giving a separate identity with logged authorisation and on least privilege as runtime control. Anyone who wants to go deeper into the control of agents will find more in our overview of agentic AI and the control of AI agents.
The news value remains the draft standard and the substantive shift it signals: from mutable observability to auditable execution evidence. But as long as the document is a working draft, it is a direction, not a seal of approval. Whoever adopts it now does so as their own governance choice, not as compliance with an established standard.
Sources and references
- Agent Audit Trail: A Standard Logging Format for Autonomous AI Systems, draft-sharif-agent-audit-trail-05
- From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents
- Reasoning Provenance for Autonomous AI Agents: Structured Behavioral Analytics Beyond State Checkpoints and Execution Traces
- NovaFabric: Tamper-Evident, Replayable Evidence for Autonomous AI Agent Runs
Sources: The article relies on the IETF Internet-Draft draft-sharif-agent-audit-trail-05 and on three arXiv publications about evidence tracing, reasoning provenance and tamper-evident replayability of AI agent runs.