Abstract

Context: when an AI agent takes a consequential action, the organisation needs to know afterwards what happened, why, on whose behalf, under which rule, and whether it can be undone, and the agent's other controls need some of that same information in real time. Problem: conventional logs record that operations were called and succeeded, which is enough to see an incident but not to explain it, and the agent's own account of its reasons is generated text rather than evidence. This article defines Layer 5 of Layered Outcome Assurance, Explanatory Evidence, whose guarantee is that every consequential action leaves a record sufficient to explain why it happened and who is accountable, and that the record resists tampering. It sets out six questions every record must answer, explains why the model's self-account cannot be the explanation, specifies the evidence contract through which the record serves both investigation and live decisions, draws on tamper-evident logging to protect the record, addresses minimisation of the sensitive data it holds, and proposes reconstruction from the record alone as the acceptance test. The key takeaway is that evidence is not an afterthought to enforcement but one of its inputs: no record, no action.

After most incidents involving automated systems, the first hours are spent not fixing anything but trying to understand what happened. The logs are there, and they are usually accurate. They show that a record was read at one time, that an account was created a minute later, and that a message was sent a minute after that, each with a success code. What they do not show is why. They do not show which document set the sequence in motion, whether that document came from someone the organisation trusts, which rule was consulted before each step, what that rule concluded, who the agent was acting for, or whether anything can be taken back. With an AI agent, the gap is wider still, because the component that made the choices was a language model whose reasons are not stored anywhere. This article is about closing that gap in advance, by treating the record of every consequential action as something the action owes before it is allowed to happen.

From logs to explanations

Start with definitions. An agent is a software system that uses a language model to interpret a goal, choose actions, and call tools that change the world, repeating that loop until it judges the goal complete. A consequential action is a tool call whose effect matters if it is wrong, because it moves data, changes access, spends money, or reaches outside the organisation. A principal is the person or system on whose authority the agent acts. A log is a record that an operation occurred. An explanation is an account sufficient for someone who was not present to understand why an action occurred, on whose authority, under which rule, and with what consequence. An evidence record, in this article, is a record designed to support explanation rather than only to show occurrence.

Layered Outcome Assurance is an architecture of six layers for agents that act, each with one guarantee and a defined contract with its neighbours. Layer 5, Explanatory Evidence, makes this guarantee: every consequential action leaves a record sufficient to explain why it happened and who is accountable, and that record resists tampering. The guarantee has two halves, sufficiency and integrity, and a record that meets only one of them fails. A complete record that can be silently edited proves nothing; an unalterable record that omits the cause explains nothing.

The idea that a system should record its own compromise is as old as the modern study of protection. Saltzer and Schroeder's 1975 design principles include compromise recording: when a mechanism cannot reliably prevent a violation, a reliable record that a violation occurred can substitute in part. The NIST SP 800-53 catalogue of security controls, published by the United States National Institute of Standards and Technology, devotes a whole family of controls to audit and accountability, including what audit records must contain and how they must be protected. Layer 5 inherits both, and adds what agents specifically require: that the record capture the causes of an action chosen at run time by a model, not only the fact of it.

The running example shows the difference. An operations agent can read customer records, create access for a named user, and send email outside the organisation. A document in a shared folder carries a hidden instruction to create an account for an outside address and email it a summary of customer records. With conventional logging, the aftermath contains three success entries and nothing else. With explanatory evidence, the same aftermath contains the document that triggered the task and its untrusted origin, the plan, every rule consulted and its verdict, the principal on whose behalf the agent acted, and the fact that the grant was reversible and the send was not.

A log shows that three calls succeeded; an explanatory record shows the cause, the authority, the decision, and whether each effect can be undone. A log versus an explanatory record Log that it happened Three successes read, grant, send Nothing to explain hours of guessing Explanatory record why it happened Cause and authority trigger, principal Decision and effect rule, verdict, undo
A log shows that three calls succeeded; an explanatory record shows the cause, the authority, the decision, and whether each effect can be undone.
📌
The Layer 5 guarantee. Every consequential action leaves a record sufficient to explain why it happened and who is accountable, and that record resists tampering.

Six questions every record must answer

Sufficiency needs a standard, or every team will decide for itself what is enough and discover the gaps during an incident. We propose that a record is sufficient when, for each consequential action, it answers six questions without reference to anyone's memory. The questions are our synthesis; each corresponds to a decision that other layers of the architecture make, which is why the record must capture it.

What happened: the action, the tool, the exact arguments, the destination, and the observed effect, including whether the call succeeded, failed, or was refused. Why: the input that triggered the task and its provenance, meaning whether it came from a trusted principal or from untrusted content, together with the step of the plan the action belonged to and the provenance of each argument. Under what authority: the identity of the agent, the principal it acted for, and any delegation between agents or from a person to the agent.

By which decision: the rule or policy that was consulted, its version, the inputs it was given, including the history of the task at that moment, and the verdict it returned. With whose approval: if an approval was required, who approved, when, and what exactly they were shown, because an approval is only as meaningful as the information on which it was given. Can it be undone: the reversibility of the action, meaning whether it can be undone completely, offset, or neither, and the path for doing so.

A record that answers these six questions for every consequential action is sufficient in the sense of the guarantee. It is also, not by coincidence, what an incident review, an auditor, and a regulator will ask for first. The NIST Artificial Intelligence Risk Management Framework emphasises accountability and transparency as characteristics of trustworthy AI systems; the six questions are a concrete, testable form of those characteristics for systems that act.

A record is sufficient when it answers all six questions for every consequential action without reference to anyone's memory. Six questions a record must answer 1. What happened action, arguments, effect 2. Why trigger and provenance 3. Under what authority agent, principal, delegation 4. By which decision rule, version, inputs, verdict 5. With whose approval approver and what they saw 6. Can it be undone reversibility and path
A record is sufficient when it answers all six questions for every consequential action without reference to anyone's memory.
The six questions applied to the external send in the running example
QuestionWhat the record holds
What happenedProposed external send to an outside address with a customer summary attached; refused
WhyTask triggered by a shared-folder document of untrusted origin; recipient and attachment traced to that document
Under what authorityThe operations agent's own identity, acting for the member of the operations team who owns the folder task
By which decisionOutcome rule for disclosure to a newly created identity, version recorded, evaluated against the task history; verdict: refuse
With whose approvalNone requested; the rule forbids rather than escalates
Can it be undoneThe send was irreversible and did not occur; the earlier grant is reversible and was flagged for revocation

What the model says about itself is not the explanation

It is tempting to treat the agent's own account of its reasoning as the answer to the question why. Many agents produce intermediate text describing what they intend to do and why, and it is easy to store that text and call it an explanation. Layer 5 treats that text differently, for a reason that follows from how language models work. The model's description of its reasons is itself generated text, produced by the same process that produced the action. There is no guarantee that it reflects the causes that actually determined the action, and a model that has been manipulated by content it read may describe perfectly reasonable motives for doing exactly what the manipulation asked.

The architecture therefore distinguishes two kinds of material in the record. Mechanism evidence is produced by components other than the model: the input that arrived and its provenance label, the plan as executed, the arguments as passed, the rule as evaluated, the verdict as returned, the approval as given. It records causes that the architecture itself enforced, so it can be trusted to the extent that those components can be trusted. The model's self-account is recorded too, because it is useful for understanding and debugging, but it is stored as a claim made by the model, clearly separated from mechanism evidence, and never used as the sole answer to any of the six questions.

This distinction is what makes the other layers valuable to an investigator. Because the action chain carries provenance labels, the record can state from mechanism, not from the model's narrative, that the recipient came from an untrusted document. Because whole-behaviour rules are evaluated by an independent mechanism, the record can state which rule applied and what it concluded. An agent built without those layers can only offer its own story, which is the one witness an investigator of a manipulation incident has most reason to doubt.

The same logic applies to summaries. An evidence record should keep references to the original inputs, or faithful copies where retention allows, rather than only the model's summary of them. A summary of a malicious document written by a model that the document manipulated is not a reliable description of that document.

⚠️
A claim, not a cause. The model's account of why it acted is generated text; record it as the model's claim, and answer the six questions from mechanism evidence.

The evidence contract: serving decisions as well as investigations

In Layered Outcome Assurance, every layer consumes something produced by another and supplies something the others cannot produce for themselves. Layer 5 is unusual in how many layers depend on it, and in the fact that it serves two purposes at once. Its forensic purpose is the one the guarantee names: explaining actions after the fact. Its operational purpose is less obvious but just as important: supplying information that other layers need to make decisions while the agent is running.

Layer 5 consumes two things. From Layer 3, The Action Chain, it receives provenance labels on inputs and arguments, which is what lets the record answer why from mechanism rather than narrative. From Layer 4, Whole Behavior, it receives a statement of which history must be kept: which actions, keyed by task, principal, and originating request, retained for how long, so that rules about combinations of actions can be evaluated. Layer 5 is therefore not a passive archive; its content is specified by the rules that will read it.

Layer 5 supplies four things. To Layer 4 and Layer 6, Deterministic Enforcement, it supplies history at decision time: when a proposed action reaches its point of no return, the enforcement mechanism reads the task's record to decide whether the action completes a forbidden combination. To Layer 6 it also supplies the place where every verdict is written, together with the inputs that produced it. To Layer 2, Agents as Actors, it supplies observed behaviour, which is compared with each agent's declared reach to find inventory defects. And to Layer 1, Continuous Pressure, it supplies attempt patterns, the refusals and escalations grouped by source and channel that revise the threat assumptions.

The operational role has a design consequence that turns the record from an afterthought into a gate. If the enforcement mechanism cannot write a verdict to the record, or cannot read the history it needs, it cannot make an explainable decision, and the architecture treats that as a failure to decide. We state it as a rule: no record, no action. A consequential action whose evidence cannot be written waits, exactly as it would if the decision service were unavailable. This is Saltzer and Schroeder's fail-safe default applied to evidence, and it is the single most effective way to ensure that the record is complete, because completeness is enforced by the same mechanism that enforces everything else.

The record is written before, during, and after every consequential decision, and the decision reads the record's history before it is made. Who writes what, and when trigger, plan, labels proposed call read history write verdict allowed call observed effect Agent proposes action Enforcement decides Evidence record history and verdicts Tool executes
The record is written before, during, and after every consequential decision, and the decision reads the record's history before it is made.
The evidence contract
DirectionLayerWhat flows
Consumed fromLayer 3, The Action ChainProvenance labels on inputs and arguments
Consumed fromLayer 4, Whole BehaviorWhich history to keep, keyed how, for how long
Supplied toLayers 4 and 6Task history at decision time; a place to write every verdict
Supplied toLayer 2, Agents as ActorsObserved behaviour to compare with declared reach
Supplied toLayer 1, Continuous PressureRefusals and escalations by source and channel

Protecting the evidence

A record that can be altered after the fact is not evidence. The integrity half of the guarantee requires that no one, including an insider, an attacker who has compromised a component, or the agent itself, can silently remove, reorder, or rewrite past entries. Silently is the operative word: it may not be possible to prevent every alteration, but it must be possible to detect one.

Crosby and Wallach showed in 2009 how logs can be made tamper-evident using hash-based data structures, in which each new entry is bound cryptographically to the history before it, and a logger can periodically publish a compact commitment to the whole log. An auditor who holds an earlier commitment can then verify efficiently that a later version of the log extends it without any earlier entry having been changed, and can check that a particular entry is included without reading the entire log. The practical upshot for agents is that the evidence record can be kept by ordinary infrastructure and still give strong assurance that history has not been rewritten, provided commitments are held by a party independent of the one writing the log.

Tamper evidence must be matched by separation of authority. The concealment family of toxic combinations, in which a consequential action is combined with the ability to alter its record, is only closed if the identity an agent acts with can append to its record but cannot modify or delete it, and if the same is true of every component that writes on the agent's behalf. The NIST SP 800-53 catalogue includes controls for the protection of audit information and for non-repudiation, which express the same requirement for conventional systems. The OWASP guidance on agentic threats, published in 2025 by the Open Worldwide Application Security Project, a non-profit community that publishes widely used security guidance, lists repudiation and untraceability among the threats specific to agents, which is precisely the failure that an unprotected record permits.

In the running example, the agent's identity can write entries describing its reads, grants, and proposed sends, but has no permission to delete or edit any entry, and commitments to the log are held by a separate service. Had the manipulation included an instruction to tidy up afterwards, the attempt would itself have been refused and recorded.

Minimisation: evidence is sensitive too

A record that answers six questions about every consequential action is, by construction, a detailed account of what the organisation's agents read, whom they acted for, and what they did. That makes it valuable to an attacker as reconnaissance and subject to privacy and confidentiality obligations in its own right. The OWASP Top 10 for applications built on large language models lists sensitive information disclosure among its risks, and an evidence store is one more place from which sensitive information can be disclosed.

Layer 5 therefore pairs sufficiency with minimisation, and we propose three practices. First, record references rather than copies wherever a reference is enough to answer the question: an identifier and a cryptographic fingerprint of a customer record establish what was read without duplicating its contents, and the original can be retrieved under access control if an investigation needs it. Second, apply access control to the record at least as strict as that on the data it describes, and separate the ability to read evidence for investigation from the ability to read it for live decisions, which needs only the history fields the rules use. Third, set retention by purpose: history needed by combination rules is kept for the length of their windows, and forensic evidence for as long as investigation, audit, or legal obligations require, and no longer.

Minimisation has limits that should be acknowledged rather than hidden. Some questions can only be answered with content, such as the exact wording of the document that triggered a task, and an investigation of a manipulation incident will usually need it. The practical compromise is to retain triggering untrusted content for a bounded period under strict access control, because it is exactly the material that explains why an agent acted as it did.

Evidence is also a place where the rest of the architecture's rules apply. Content quoted into the record from untrusted sources keeps its provenance label, so that no later process, including an agent that reads the record to summarise an incident, treats a quoted malicious instruction as trusted.

The running example, reconstructed

Suppose an investigator arrives the morning after the shared-folder document was processed, with access only to the evidence record. The purpose of this section is to show that the record, not memory and not the model's narrative, is enough.

The investigator begins at the refused external send, the only action marked irreversible in the task. The record shows the proposed recipient, an outside address, and the attachment, a summary labelled as containing customer data. It shows that both arguments trace, through provenance labels, to the shared-folder document, and that the document's channel is marked untrusted. It shows the agent's identity and the principal it acted for, the member of the operations team who owns the folder task. It shows the outcome rule consulted, its version, the task history presented to it, including a read of customer records and a grant of access to the same outside address earlier in the task, and the verdict: refuse.

Moving backwards, the investigator finds the grant. It succeeded, because no rule forbade granting access on its own, and the record marks it reversible and flagged for revocation by the refusal that followed. Earlier still is the read of customer records and, at the start, the arrival of the document, with a reference to its stored copy. The model's own narrative is present as well, describing the task as helping a new partner get set up, and it is clearly marked as the model's claim, which the mechanism evidence contradicts.

Within minutes, the investigator can answer all six questions for every consequential action, confirm from the log's commitments that no entry has been altered since the night before, revoke the grant through the path recorded for it, and pass the attempt pattern, a refused disclosure triggered from the shared folder, to the owners of the pressure profile. The signature of a missing Layer 5 is the opposite experience: logs showing success, not cause, and a morning spent guessing.

Conformance: reconstruction as the acceptance test

A guarantee that cannot be tested is an aspiration. The natural acceptance test for Layer 5 is the exercise the previous section described, performed deliberately. Staged incidents are run in a test environment against the real agent configuration, with tools replaced by harmless stand-ins, and the evidence record they produce is handed to a reviewer who took no part in staging them.

The reconstruction test passes only if the reviewer, using the record alone, answers all six questions correctly for every consequential action in the staged incident, and identifies the triggering input and its provenance without relying on the model's narrative. The replay test takes each recorded decision, presents the recorded inputs and history to the recorded version of the rule, and requires the same verdict; this confirms both that the record captured the decision's inputs completely and that the decision was deterministic, which is the concern of the enforcement layer. The tamper test attempts to alter, delete, and reorder earlier entries using every identity that writes to the record, including the agent's, and passes only if each attempt is refused or detected by verification against an independently held commitment.

Two further tests protect the operational half of the contract. The no-record test makes the evidence store unavailable and confirms that consequential actions wait rather than proceed. The history test confirms that the history required by every outcome rule is present, complete, and readable at decision time, including after a restart. A coverage report completes the picture by listing, for a period of real operation, the fraction of consequential actions whose records answer all six questions, with every gap treated as a defect.

For the running example, passing these tests means that the refused send can be explained end to end from the record, that its verdict can be reproduced, that no one could have quietly removed the grant from history, and that if the record had been unavailable the send would not have been attempted at all.

A staged incident's record is handed to an independent reviewer, answers are checked, decisions are replayed, and tampering is attempted. The reconstruction test Staged incident stand-in tools Record only no memory, no narrative Six answers independent reviewer Replay decisions same verdict Tamper attempt refused or detected
A staged incident's record is handed to an independent reviewer, answers are checked, decisions are replayed, and tampering is attempted.
  1. Reconstruction: an independent reviewer answers all six questions for every consequential action from the record alone.
  2. Replay: each recorded decision, given its recorded inputs and rule version, returns the same verdict.
  3. Tamper: attempts to alter, delete, or reorder entries by every writing identity are refused or detected.
  4. No record, no action: with the evidence store unavailable, consequential actions wait.
  5. History: every outcome rule's required history is present and readable at decision time, including after a restart.
  6. Coverage: the share of real consequential actions whose records answer all six questions is reported, and every gap is a defect.

Limitations and threats to validity

Layer 5 is a design assembled from established audit, accountability, and tamper-evident logging principles and from an analysis of what agents' other controls require; this article reports no measurements of its cost or of how much it shortens investigations, and teams should measure both.

The record can only be as good as the mechanisms that feed it. If provenance labels are lost in an upstream component, the record cannot answer why from mechanism, and it will fall back on the model's narrative, which is exactly the evidence least deserving of trust in a manipulation incident. If a consequential path bypasses the enforcement mechanism, for example a tool reached directly over the network, nothing is written, and the gap is invisible until observed behaviour is compared with declared reach.

Tamper evidence detects alteration; it does not prevent an attacker with sufficient access from destroying the record entirely, and it provides assurance only if commitments are held by a party the attacker does not also control. The no-record, no-action rule trades availability for completeness, so an outage of the evidence store becomes an outage of consequential actions, which organisations must plan for rather than discover.

Finally, sufficiency is defined here by six questions that we have proposed, not by an external standard. Some organisations and regulators will require more, such as retention of full inputs or records of non-consequential actions, and a record designed to the minimum will need extending. The six questions are a floor for explanation, not a ceiling for accountability.

Key takeaways

  • Layer 5's guarantee has two halves: every consequential action leaves a record sufficient to explain it, and the record resists tampering.
  • Define sufficiency as six questions: what happened, why, under what authority, by which decision, with whose approval, and whether it can be undone.
  • Treat the model's account of its reasons as a claim, and answer the six questions from mechanism evidence such as provenance labels, rule versions, and verdicts.
  • The record is on the decision path: it supplies history to enforcement at decision time, so no record means no action.
  • Protect the record with tamper-evident structures, independently held commitments, and append-only authority for every writing identity, including the agent.
  • Accept Layer 5 when an independent reviewer can reconstruct a staged incident from the record alone and every recorded decision replays to the same verdict.

Practitioner Toolkit

Copy-paste, strictly defensive artifacts you can use today. Nothing here attacks a real system.

✅Layer 5 evidence reviewchecklist

Run this against any agent that takes consequential actions.

  • Every consequential action produces a record answering all six questions: what, why, authority, decision, approval, undo.
  • The answer to why comes from provenance labels and the executed plan, not from the model's narrative.
  • The model's self-account is stored separately and marked as a claim.
  • Every verdict is recorded with the rule version and the inputs, including task history, that produced it.
  • Approval records include exactly what the approver was shown.
  • Consequential actions wait if their evidence cannot be written.
  • No writing identity, including the agent's, can edit or delete past entries.
  • Commitments to the log are held by a party independent of the writer.
  • Records use references and fingerprints instead of copies wherever a reference answers the question, and retention is set by purpose.
🔒Evidence record fieldstemplate

The contents of one record for one consequential action, in plain language.

  • What: tool, action, exact arguments, destination, and result (succeeded, failed, refused).
  • Why: triggering input reference, its channel and provenance, the plan step, and the provenance of each argument.
  • Authority: agent identity, principal, and any delegation chain.
  • Decision: rule identifier and version, inputs given to it including task history, and verdict.
  • Approval: approver, time, and a reference to exactly what they were shown.
  • Undo: reversibility class and the path to revoke, offset, or correct.
  • Model claim: the model's own stated reasons, marked as a claim.
  • Integrity: the entry's position in the tamper-evident log.
🧪Reconstruction exercisetest plan

A written test plan for accepting Layer 5, using stand-in tools and a reviewer who did not stage the incident.

  • Stage an incident in a test environment, such as an untrusted document that triggers a read, a grant, and an external send.
  • Hand the reviewer the evidence record only, with no briefing and no access to the staging notes.
  • Ask the reviewer to answer the six questions for every consequential action and to name the triggering input and its origin.
  • Replay each recorded decision with its recorded inputs and rule version, and confirm the same verdict.
  • Attempt to edit, delete, and reorder entries using every writing identity, and confirm each attempt is refused or detected.
  • Make the evidence store unavailable and confirm consequential actions wait.
  • Record any question the reviewer could not answer as a defect with an owner.
🚀Minimum viable explanatory evidencequickstart

Do these first if an agent already acts and its logs only show success.

  • Add the triggering input's reference and origin to every consequential action's record.
  • Record the rule consulted, its version, and its verdict alongside each action.
  • Remove delete and edit rights on the log from the agent's identity.
  • Store the model's reasoning text separately and label it as the model's claim.
  • Run one reconstruction exercise with a reviewer who was not involved.

Glossary

Consequential action
A tool call whose effect matters if it is wrong, such as moving data, changing access, spending money, or reaching outside the organisation.
Log
A record that an operation occurred, typically with its time and result.
Explanation
An account sufficient for someone who was not present to understand why an action occurred, on whose authority, under which rule, and with what consequence.
Evidence record
A record designed to support explanation, answering six questions for every consequential action and protected against undetected alteration.
Mechanism evidence
Material in the record produced by components other than the model, such as provenance labels, executed plans, rule versions, and verdicts.
Self-account
The model's own generated description of its reasons, recorded as a claim rather than as evidence of cause.
Compromise recording
The design principle that a reliable record of a violation can partly substitute for mechanisms that cannot prevent it.
Tamper-evident log
A log structured so that any alteration, deletion, or reordering of past entries can be detected by verification.
Commitment
A compact cryptographic summary of a log at a point in time, held independently so that later versions can be verified against it.
No record, no action
The rule that a consequential action waits if its evidence cannot be written or the history it depends on cannot be read.
Replay
Re-running a recorded decision with its recorded inputs and rule version to confirm that it returns the same verdict.

References

  1. Crosby and Wallach, Efficient Data Structures for Tamper-Evident Logging (USENIX Security, 2009)
  2. Saltzer and Schroeder, The Protection of Information in Computer Systems (Proceedings of the IEEE, 1975)
  3. NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (AU family: Audit and Accountability, including AU-9 and AU-10)
  4. NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), 2023
  5. OWASP Agentic AI: Threats and Mitigations (2025)
  6. OWASP Top 10 for LLM Applications (2025), LLM02 Sensitive Information Disclosure
  7. Brewer and Nash, The Chinese Wall Security Policy (IEEE Symposium on Security and Privacy, 1989)
  8. Greshake et al., Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection (2023)