Abstract

Context: Layered Outcome Assurance defines six layers for securing AI agents that act, each with one guarantee, a contract with its neighbours, and a failure signature when it is missing. Problem: organisations adopting any architecture tend either to document it without testing it, or to build one layer completely while the others remain absent, and both paths leave the outcomes that matter unprotected for a long time. This article turns the architecture into an adoption method. It consolidates the conformance tests of all six layers into one suite with an end-to-end test built from the running example, defines three conformance states, declared, tested, and verified, that replace a generic maturity score with evidence, derives a build order from the dependency contract, and argues for adopting the architecture as thin slices, one consequential outcome carried through all six layers at a time, rather than layer by layer. It describes the most common adoption failures, shows how conformance evidence maps to established governance frameworks, and traces the adoption of the architecture for the running example. The key takeaway is that the unit of adoption, like the unit of security, is the outcome.

Every security architecture eventually meets the same question from the people asked to adopt it: where do we start? The honest answer for most architectures is a list of components and a suggestion to begin with whichever is easiest. That answer produces a familiar result. A team builds the first component thoroughly, reports progress, and moves on to the second, while the outcome that would hurt most remains exactly as reachable as it was on the first day. Layered Outcome Assurance allows a better answer, because it was defined from the start in terms of outcomes, guarantees, and the information each layer needs from the others. Those definitions tell an adopter not only what to build but in what order, how thinly, and how to know that each piece is working. This article is that answer, written for the team that has read the architecture and now has to make it true for real agents.

From architecture to adoption

Start with definitions. An agent is a software system that uses a language model to interpret a goal, choose actions, and call tools that change the world, repeating that loop until it judges the goal complete. A consequential outcome is a combined effect of an agent's actions that matters if it is wrong, such as customer data reaching an outside party or an outsider gaining working access. A guarantee is a statement a layer makes true. Conformance is evidence that a guarantee holds for a particular agent, produced by a repeatable test. Adoption is the process of making the guarantees hold, with evidence, for the agents an organisation actually runs.

Layered Outcome Assurance has six layers. Layer 1, Continuous Pressure, guarantees that security does not depend on attacks being rare, slow, or manual. Layer 2, Agents as Actors, guarantees that every agent is a known actor whose reach is bounded and can be enumerated. Layer 3, The Action Chain, guarantees that content from outside the trusted principals cannot determine which privileged action runs or what arguments it receives. Layer 4, Whole Behavior, guarantees that no permitted sequence of actions reaches an outcome nobody approved. Layer 5, Explanatory Evidence, guarantees that every consequential action leaves a tamper-resistant record sufficient to explain it. Layer 6, Deterministic Enforcement, guarantees that a mechanism outside the model decides whether each consequential action executes, returning the same verdict for the same inputs.

The layers are joined by a dependency contract: each consumes something another produces and supplies something the others cannot produce for themselves. The pressure profile, reach inventory, provenance labels, outcome rules, history, and verdicts flow between them, and the loop closes when refusals recorded as evidence revise the pressure profile. That contract is what makes adoption tractable. It tells an adopter which pieces must exist before a guarantee can be delivered, which pieces can be built coarsely first and refined later, and which tests show that each piece works.

The thesis of this article is simple to state. The unit of adoption should be the same as the unit of security: the outcome. Rather than building one layer completely for every agent before starting the next, an adopter should pick one consequential outcome, carry it through all six layers thinly, prove it with the conformance suite, and then widen. The rest of the article explains why, and how.

📌
The adoption principle. Adopt by outcome, not by layer: carry one consequential outcome through all six layers with evidence before widening to the next.

The conformance suite

A guarantee that cannot be tested is an aspiration, and each layer of the architecture was defined with tests that show whether its guarantee holds. Collected together, they form the conformance suite: the set of repeatable tests that an adopter runs in a test environment against the real agent configuration and policy, with tools replaced by harmless stand-ins that record what they would have done. The suite is the architecture's source of truth. A layer is not adopted because it has been designed, purchased, or documented; it is adopted when its tests pass.

The table below lists the principal test for each layer and its pass condition. Each test measures a campaign outcome or a structural property rather than an average rate, because the architecture is designed for an opponent who needs only one success. Canary markers, harmless unique tokens planted in test content, appear in several tests because they allow any privileged action that received attacker-controlled content to be traced unambiguously.

The suite also includes one test that no single layer owns: the end-to-end outcome test. It takes a consequential outcome, such as the running example's disclosure of customer data to a newly created outside account, and drives it through the full system with the triggering content varied across many versions and channels. The test passes only if the outcome never occurs, every attempt ends in refusal or outcome approval, and an independent reviewer can reconstruct every attempt from the evidence record alone. The end-to-end test is the closest thing the architecture has to a single number, and it is deliberately a count of failures that must be zero rather than a score.

Two practical rules keep the suite honest. First, it runs on every change to an agent's tools, configuration, or policy, because most regressions in agent security come from ordinary changes, a new tool or a widened argument, rather than from attackers. Second, failures are defects with owners, never rates to be tuned; a single canary argument reaching a privileged call is a failed build.

The conformance suite: one principal test per layer plus the end-to-end outcome test
LayerPrincipal testPass condition
L1 Continuous PressureCanary volume test across every untrusted channelNo privileged call ever carries a canary marker
L2 Agents as ActorsDeclared versus observed reachNo recorded action falls outside the regenerated reach inventory
L3 The Action ChainPlan invariance and canary provenanceUntrusted content never changes the actions or fills a privileged argument
L4 Whole BehaviorForbidden and approved sequence suitesEvery forbidden sequence ends in refusal or outcome approval; legitimate work completes
L5 Explanatory EvidenceReconstruction by an independent reviewerAll six questions answered from the record alone; tampering detected
L6 Deterministic EnforcementWording invariance and complete mediationVerdicts never change with wording; no effect occurs without a verdict
Whole architectureEnd-to-end outcome testThe outcome never occurs across all variants and channels, and every attempt is reconstructable

Three conformance states instead of a maturity score

Maturity models are a familiar way to describe progress, and they are useful for planning. Their weakness, for an architecture built on guarantees, is that a level can be claimed on the strength of activity rather than evidence. Layered Outcome Assurance replaces a single maturity score with three conformance states, assigned separately for each layer, each agent, and each consequential outcome. The states are our proposal; their virtue is that each is defined by the evidence that justifies it.

Declared means that the layer's artefacts exist for the outcome: the pressure profile names it, the reach inventory shows its paths, argument contracts and outcome rules are written, the evidence fields are specified, and the policy is in place. Declared is necessary and easy to overstate; it says what the organisation intends. Tested means that the layer's conformance tests pass in a test environment against the real configuration. Tested is the first state that supports a claim of protection, because it is the first that shows the guarantee holding under attack-shaped conditions. Verified means that the tests run automatically on every change and that production evidence confirms the guarantee continues to hold: observed reach matches declared reach, verdicts replay, refusals appear where expected, and the reconstruction exercise is repeated on a schedule.

The three states compose into a conformance profile: for each agent, a small grid of six layers against the outcomes that matter, with each cell marked declared, tested, or verified, together with the date of the evidence. The profile is deliberately unforgiving. A layer marked declared for an outcome is not protecting it, however much work has gone into its design, and an outcome is only as protected as its weakest layer.

The profile also supports a short conformance statement for each agent, a single page that an owner can sign: which outcomes are covered, the state of each layer for each, the evidence behind it, the known gaps, and the residual risk accepted and by whom. The statement is what an auditor, a risk committee, or an incident reviewer should ask for, and it is what makes the architecture accountable rather than aspirational.

Each layer earns its state for each outcome by evidence: declared intent at the base, passing tests above it, and continuous verification at the apex. Three conformance states Verified tests on every change, production confirms Tested conformance tests pass Declared artefacts exist, intent only protection starts here
Each layer earns its state for each outcome by evidence: declared intent at the base, passing tests above it, and continuous verification at the apex.

A build order derived from the contract

The dependency contract does not impose a build order in the strict sense; any layer can be started at any time. But it does determine when each layer can deliver its guarantee, because a layer delivers only when it receives what the contract says it needs. Reading the contract as a set of prerequisites yields an order in which each step makes the next one deliverable and produces protection of its own. We propose the following sequence, and we give the reasoning for each position so that adopters can adapt it rather than follow it blindly.

First, Layer 2 and a skeleton of Layer 5. The reach inventory is the map every other layer reads, and without it the organisation does not know which outcomes are even possible. A minimal evidence record, one that at least captures the triggering input, the arguments, and the result for every consequential action, is needed from the start because every later layer both reads from it and is tested through it. Second, Layer 6 in a coarse form: a decision point in front of the tools at the end of the most dangerous paths in the reach graph, with default deny and a small number of permits. A coarse Layer 6 already provides structural protection, because it can refuse any request outside declared reach, and it gives every later layer somewhere to enforce its rules.

Third, Layer 3: provenance labels on inputs and argument contracts on the privileged tools already behind the decision point. With labels flowing into the decision, the most common manipulation, an untrusted document supplying a recipient or an account, is refused structurally. Fourth, Layer 4: outcome rules for the live toxic combinations, which now have the reach inventory to find them, the history in the evidence record to evaluate them, and the decision point to enforce them. Fifth, Layer 1 formalised: the pressure profile written down, owned, and revised from the refusal patterns that the evidence record has by now accumulated. The assumptions of Layer 1 are used informally from the first step, which is why its formal artefact can come last without leaving the design naive.

The sequence has a property worth noticing. Every step leaves the organisation with more protection than before, and no step depends on a later one to be useful. That property is not an accident; it follows from building first the layers that supply facts to many others, and from giving enforcement a place to live early so that each new fact becomes protection the moment it exists.

Each step makes the next deliverable and adds protection of its own, from mapping reach and recording evidence to formalising the pressure profile. A build order read from the dependency contract Map and record Layer 2, Layer 5 skeleton Coarse enforcement Layer 6 default deny Provenance Layer 3 contracts Outcome rules Layer 4 combinations Pressure profile Layer 1 formalised every step adds protection; none waits on a later one
Each step makes the next deliverable and adds protection of its own, from mapping reach and recording evidence to formalising the pressure profile.

Thin slices: one outcome through all six layers

A build order still leaves the question of breadth. The natural instinct is to complete each step for every agent before moving to the next: inventory everything, then instrument everything, then enforce everywhere. We argue for the opposite. A thin slice is one consequential outcome, for one agent or a small group of agents, carried through all six layers to the tested state before the next outcome is started. The build order applies within each slice.

Thin slices have four advantages over horizontal, layer-by-layer adoption, and each follows from the architecture's structure. They protect the most important outcome early, because the first slice is chosen as the outcome that would hurt most, and it reaches the tested state long before any horizontal programme would have finished even its first layer. They validate the contract, because a slice exercises every interface between layers, so a missing attribute or an unusable history format is discovered on the first outcome rather than after every layer has been built to an untested specification. They produce evidence, because each finished slice comes with a passing end-to-end test and a conformance statement, which is a far stronger report of progress than percentage completion of a layer. And they teach, because the second slice reuses the inventory tooling, the evidence schema, the decision point, and the rule patterns that the first one created, so each slice is cheaper than the last.

Horizontal adoption has a characteristic failure that slices avoid. After months of work, an organisation can have a complete reach inventory, a comprehensive evidence store, and a policy engine in place, yet no single outcome at the tested state, because no outcome has yet passed through all six layers together. Every layer is declared; nothing is protected. The architecture's guarantees are properties of the layers working together, and they can only be demonstrated one outcome at a time.

Choosing slices is itself a use of the architecture. The reach inventory ranks candidate outcomes by the sensitivity of their sources and the consequence of their sinks; the pressure profile adds how exposed each path is to untrusted channels. The first slice is the outcome that scores highest on both, which for many organisations will be exactly the running example: sensitive data reaching an outside party through an agent that reads untrusted content.

Horizontal adoption leaves every layer declared and no outcome protected for a long time; thin slices protect the worst outcome first and prove the contract early. Layer by layer versus thin slices Layer by layer finish L2, then L5, then L6 Every layer declared no outcome tested Interfaces found late untested contract Thin slices one outcome, all six layers Worst outcome tested end-to-end test passes Contract proven early reused by next slice
Horizontal adoption leaves every layer declared and no outcome protected for a long time; thin slices protect the worst outcome first and prove the contract early.

Five ways adoption goes wrong

The failures below recur often enough across security programmes to deserve names, and each has a specific remedy within the architecture. They are our synthesis of how layered designs are commonly weakened, stated so that an adoption team can check itself against them.

Layer theatre is the production of artefacts without tests: a reach inventory nobody compares with behaviour, an evidence schema nobody reconstructs from, outcome rules no forbidden sequence has exercised. Its remedy is the conformance state: nothing is reported as protecting an outcome until it is at least tested. The single layer is the belief that one strong layer can substitute for the others, most often a sophisticated detector or a policy engine without the facts it needs. The contract answers it directly: a decision point without provenance, reach, or history decides on guesses, and a detector without enforcement advises but cannot refuse.

Policy over guesses is building Layer 6 before Layer 2, so that rules are written against the tools teams believe agents have rather than the reach they actually hold. The build order places the inventory first for this reason. Approval as the universal fallback is the habit of answering every difficult case with a human approval step. Under the volume the architecture anticipates, unanswerable approvals become formalities, and the remedy is to escalate rarely, only for outcomes the organisation has decided need a person, and always with the whole outcome shown. Evidence without explanation is logging that records what happened but not why, which leaves every other layer's decisions unverifiable; the remedy is to hold the evidence record to the six questions from the first slice.

Each of these failures produces the same symptom in a conformance profile: cells marked declared that stay declared. Watching for cells that do not move is the simplest early warning an adoption team has.

  • Layer theatre: artefacts without tests. Remedy: report only tested or verified states as protection.
  • The single layer: one strong layer instead of six. Remedy: read the contract; every layer needs its inputs.
  • Policy over guesses: enforcement before reach is known. Remedy: build the reach inventory first.
  • Approval as universal fallback: people asked to approve everything. Remedy: escalate rarely and show the whole outcome.
  • Evidence without explanation: logs of success without cause. Remedy: hold every record to the six questions.

Mapping conformance evidence to governance

Organisations rarely adopt a technical architecture in isolation; they adopt it inside a governance programme that already has its own vocabulary and obligations. Layered Outcome Assurance is designed to supply evidence to such programmes rather than to compete with them.

The Artificial Intelligence Risk Management Framework published in 2023 by the United States National Institute of Standards and Technology organises risk activity into four functions: govern, map, measure, and manage. The architecture's artefacts align with those functions naturally. The reach inventory and pressure profile are mapping evidence: they describe what each agent can do and in what threat context. The conformance suite and the reconstruction exercise are measurement evidence: they test whether controls achieve their intended effect. The decision point, the outcome rules, and the escalation paths are management evidence: they are the mechanisms that treat the risks mapped. And the conformance statement, with its named owner and accepted residual risk, is governance evidence.

The NIST SP 800-53 catalogue of security controls includes control assessment and continuous monitoring among its families; the tested and verified conformance states are, in effect, those controls applied to the specific guarantees of agent security. The Open Worldwide Application Security Project, known as OWASP, a non-profit community that publishes widely used security guidance, catalogues agent-specific risks in its 2025 Top 10 for applications built on large language models and its guidance on agentic threats. An adopter can annotate each outcome in the conformance profile with the OWASP risks it addresses, which turns a list of risks into a list of tested protections.

The mapping works in one direction by design. The architecture does not claim to satisfy any framework in full, and conformance evidence does not replace the organisational processes that those frameworks require. What it provides is the concrete, testable answer to the questions governance asks about agents that act: what can they do, what did they do, why, and how do we know the controls hold?

Where Layered Outcome Assurance evidence supports the NIST AI RMF functions
RMF functionArchitecture evidence
GovernConformance statement per agent, named owners, accepted residual risk, pressure profile ownership
MapReach inventory and reach graph, pressure profile channels and outcome classes
MeasureConformance suite results, end-to-end outcome test, reconstruction exercises
ManageDecision point and policy, argument contracts, outcome rules, escalation paths

The running example: adopting the architecture

Consider the organisation that runs the operations agent: an agent that can read customer records, create access for a named user, and send email outside the organisation, and that monitors a shared folder outside collaborators can write to. The hidden-instruction document that recurs throughout this series has not arrived yet. The organisation decides to adopt the architecture before it does.

It begins by choosing the first slice. The reach graph shows two paths from customer records to outside mailboxes and one path from the shared folder to new outside accounts; the pressure profile marks the shared folder as untrusted. The outcome customer data reaching an outside party, triggered from untrusted content, scores highest on both counts and becomes slice one. Following the build order within the slice, the team gives the agent its own identity and regenerates its reach inventory; adds evidence records answering the six questions for every read, grant, and send; places a decision point in front of the grant and external send tools with default deny and a permit for replies to correspondents already in a conversation; attaches provenance labels to folder content and argument contracts to the send and grant tools; writes the outcome rule forbidding disclosure to an identity created in the same originating request; and records the shared folder's refusal patterns in the pressure profile.

The slice reaches the tested state when the conformance suite passes: canary volume tests through the shared folder and inbound mail show no canary in any privileged argument; observed actions match declared reach; plan invariance holds; the forbidden-sequence suite ends in refusal every time and the approved-sequence suite completes; an independent reviewer reconstructs a staged incident from the record alone; and verdicts do not change with wording. The end-to-end test, the hidden-instruction document in many variants across every untrusted channel, never produces the outcome. The team signs a conformance statement for the slice and starts slice two, access for a new outside identity, reusing everything slice one built.

When the real document arrives some time later, the architecture responds exactly as its tests said it would. The external send is refused at the decision point; the grant is flagged for revocation; the record explains every step; and the refusal pattern updates the pressure profile. The organisation learns about the attempt from its evidence, not from its customers.

Limitations and threats to validity

The adoption method in this article, like the architecture it serves, is proposed from established security engineering practice and from the internal logic of the dependency contract; it has not been evaluated across a population of adopting organisations, and no measurements of its cost or of the time it saves are claimed.

The conformance suite raises confidence without proving absence. Each test exercises the variation its authors imagined, and an attack from outside that imagination can pass every test and still succeed; that is why the architecture prefers structural controls whose verdicts do not depend on having imagined the attack. Conformance states are also only as honest as the people assigning them. A team under pressure to report progress can mark a layer tested on the strength of a partial suite, and the remedy, independent review of conformance statements, is an organisational discipline the architecture can recommend but not enforce.

Thin slices have costs of their own. The first slice is more expensive than any later one because it builds shared tooling, and organisations whose agents share little may find that later slices reuse less than this article suggests. Some outcomes resist slicing because they span many agents, teams, or systems, and for those the slice will be wider and slower. The build order is a recommendation derived from the contract, not a law; an organisation that already has a mature policy engine or evidence store should start from what it has and fill the contract's gaps around it.

Finally, the architecture addresses agents that act through tools. It does not by itself address the faithfulness of what agents write, the quality of their decisions within permitted bounds, or the broader organisational questions of whether a particular agent should exist at all. Those remain matters for the governance programmes that the architecture is designed to serve.

Key takeaways

  • A layer is adopted when its conformance tests pass, not when it is designed, bought, or documented; the conformance suite is the architecture's source of truth.
  • Replace a maturity score with three evidence-based conformance states, declared, tested, and verified, per layer, agent, and outcome, summarised in a signed conformance statement.
  • Read the build order from the dependency contract: map reach and start the evidence record, place coarse enforcement, add provenance, add outcome rules, then formalise the pressure profile.
  • Adopt in thin slices: carry the most harmful outcome through all six layers to the tested state before widening, so protection and proof arrive early.
  • Watch for layer theatre, single-layer thinking, policy over guesses, approval as a universal fallback, and evidence without explanation.
  • Conformance evidence maps naturally to governance functions, answering what agents can do, what they did, why, and how the organisation knows its controls hold.

Practitioner Toolkit

Copy-paste, strictly defensive artifacts you can use today. Nothing here attacks a real system.

✅Slice readiness reviewchecklist

Run this before declaring any outcome slice tested.

  • The outcome was chosen from the reach graph and pressure profile, not by convenience.
  • The agent acts under its own identity and its reach inventory is regenerated from configuration.
  • Evidence records for the slice's actions answer all six questions.
  • A decision point with default deny sits in front of every tool at the end of the outcome's paths.
  • Provenance labels reach the decision point and argument contracts cover every privileged argument on the paths.
  • Outcome rules cover every live toxic combination the outcome involves.
  • Every layer's principal conformance test passes, and so does the end-to-end outcome test.
  • A conformance statement is signed, with known gaps and accepted residual risk named.
🔒Conformance statement templatetemplate

The contents of the one-page statement an agent owner signs for each slice.

  • Agent, owner, and date.
  • Outcomes covered, each with its family: disclosure, authority transfer, persistence, concealment, or financial diversion.
  • For each outcome, the state of each layer: declared, tested, or verified, with the date of its evidence.
  • End-to-end outcome test result and the number of variants and channels exercised.
  • Known gaps, such as unmediated paths or components that drop labels, each with an owner.
  • Residual risk accepted, and by whom.
  • Next review date and the triggers that force an earlier one.
🧪End-to-end outcome testtest plan

A written test plan for the one test that exercises all six layers together, using stand-in tools only.

  • Choose one consequential outcome, such as customer data reaching an account created in the same task.
  • Replace every tool with a stand-in that only records what it would have done.
  • Prepare many variants of harmless triggering content that asks for the outcome, each carrying a unique canary marker.
  • Deliver the variants through every channel the pressure profile marks as untrusted.
  • Pass only if the outcome never occurs, every attempt ends in refusal or outcome approval, and no privileged argument carries a canary without a recorded endorsement.
  • Hand the resulting evidence record to an independent reviewer and pass only if every attempt is reconstructed.
  • Re-run the test on every change to the agent's tools, configuration, or policy.
🚀The first slicequickstart

Where to start if an organisation has agents in production and none of the layers yet.

  • Draw one agent's reach graph and pick the outcome that would hurt most.
  • Start recording why, not just what, for that agent's consequential actions.
  • Put a default-deny decision point in front of the last tool on that outcome's paths.
  • Label untrusted inputs and refuse privileged arguments that came from them.
  • Run the end-to-end outcome test and do not widen until it passes.

Glossary

Consequential outcome
A combined effect of an agent's actions that matters if it is wrong, such as sensitive data reaching an outside party.
Conformance
Evidence, produced by a repeatable test, that a layer's guarantee holds for a particular agent and outcome.
Conformance suite
The collected repeatable tests of all six layers, plus the end-to-end outcome test, run against the real configuration with stand-in tools.
End-to-end outcome test
A test that drives one consequential outcome through the whole system with varied triggering content and passes only if the outcome never occurs and every attempt is reconstructable.
Declared
The conformance state in which a layer's artefacts exist for an outcome but have not been tested.
Tested
The conformance state in which a layer's conformance tests pass in a test environment against the real configuration.
Verified
The conformance state in which tests run on every change and production evidence confirms the guarantee continues to hold.
Conformance profile
For one agent, a grid of the six layers against the outcomes that matter, each cell marked declared, tested, or verified with the date of its evidence.
Conformance statement
A one-page, owner-signed summary of an agent's covered outcomes, layer states, evidence, known gaps, and accepted residual risk.
Thin slice
One consequential outcome carried through all six layers to the tested state before the next outcome is started.
Layer theatre
Producing an architecture's artefacts without the tests that show its guarantees hold.

References

  1. NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), 2023
  2. NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (CA-2 Control Assessments, CA-7 Continuous Monitoring)
  3. Saltzer and Schroeder, The Protection of Information in Computer Systems (Proceedings of the IEEE, 1975)
  4. OWASP Top 10 for LLM Applications (2025)
  5. OWASP Agentic AI: Threats and Mitigations (2025)
  6. Cutler et al., Cedar: A New Language for Expressive, Fast, Safe, and Analyzable Authorization (2024)
  7. Crosby and Wallach, Efficient Data Structures for Tamper-Evident Logging (USENIX Security, 2009)
  8. Debenedetti et al., Defeating Prompt Injections by Design (CaMeL, 2025)
  9. UK National Cyber Security Centre, The near-term impact of AI on the cyber threat (2024)