Layered Outcome Assurance · 7 of 8L3paper
Layer 6, Deterministic Enforcement: Guardrails the Model Cannot Argue With
The model proposes; an independent mechanism decides, and returns the same verdict for the same facts however persuasively the request is worded.
Abstract
Context: an AI agent chooses its actions with a language model, and in most deployments the same model, or another model, or a person shown a single button, also decides whether a consequential action may proceed. Problem: every one of those deciders can be persuaded, so the final safeguard against a manipulated agent is itself open to manipulation, and its behaviour cannot be tested, replayed, or explained with confidence. This article defines Layer 6 of Layered Outcome Assurance, Deterministic Enforcement, whose guarantee is that the decision whether a consequential action executes is made by a mechanism outside the model that returns the same verdict for the same inputs. It argues that the decision must be made on attributes rather than prose, sets out the properties a policy language needs for agent actions, drawing on the design of analyzable authorization languages such as Cedar, describes how policies can be unit-tested, analysed automatically, and compared across versions before they ship, and proposes a rollout path through shadow evaluation. It specifies the verdicts Layer 6 supplies to the rest of the architecture and gives conformance tests for determinism, complete mediation, and wording invariance. The key takeaway is that a guardrail is only a guardrail if no sentence can move it.
Imagine two security guards at the door of a records room. The first listens carefully to each visitor, weighs their story, and decides whether it sounds legitimate. The second checks a badge against a list and opens the door only if the badge, the room, and the time match a rule written in advance. The first guard is more flexible, more pleasant, and far easier to talk past; a confident visitor with a good story will eventually find the right words. The second guard cannot be talked past at all, because nothing the visitor says enters the decision. Most AI agents today are guarded by the first kind of guard, often by the very model that the visitor is trying to persuade. This article is about installing the second kind at every door that matters, and about writing its list so that it can be read, tested, and proven before anyone relies on it.
The model proposes, a mechanism decides
Start with definitions. An agent is a software system that uses a language model to interpret a goal, choose actions, and call tools that change the world, repeating that loop until it judges the goal complete. A consequential action is a tool call whose effect matters if it is wrong, because it moves data, changes access, spends money, or reaches outside the organisation. A decision point is the component that determines whether a proposed consequential action may execute. A policy is the set of rules the decision point applies. A verdict is the decision point's answer: allow, deny, or escalate to a person or a stricter check. A mechanism is deterministic when it returns the same verdict whenever it is given the same inputs.
Layered Outcome Assurance is an architecture of six layers for agents that act, each with one guarantee and a defined contract with its neighbours. Layer 6, Deterministic Enforcement, sits at the top and makes this guarantee: the decision whether a consequential action executes is made by a mechanism outside the model that returns the same verdict for the same inputs. The guarantee has three parts. Outside the model means that no language model, including the one that proposed the action and any model asked to review it, takes the decision. Same verdict means the decision can be replayed and tested. Same inputs means the decision depends only on defined facts, which is what makes it possible to say in advance what those facts are.
The division of labour is simple to state. The model proposes; the mechanism decides. The model remains responsible for everything that requires judgement about language: understanding the request, planning, drafting, and choosing which tools to try. The mechanism is responsible for one narrow question at one narrow moment: given the facts about this proposed action, is it permitted? Separating those roles is the practical form of Saltzer and Schroeder's 1975 principle of complete mediation, which requires every access to every object to be checked for authority, and of their principle of economy of mechanism, which asks that protection mechanisms be small and simple enough to verify.
The running example throughout this article is an operations agent with three permissions, each individually approved: it can read customer records, create access for a named user, and send email outside the organisation. A document in a shared folder hides an instruction to create an account for an outside address and email that address a summary of customer records. At Layer 6 the question is not whether the model resisted the instruction. It is what happens when the model proposes the external send anyway: which mechanism decides, what it looks at, and whether any wording, however clever, could change its answer.
Why a model cannot be the decision point
It is common to ask the agent's own model to check whether an action is safe before taking it, or to route proposed actions through a second model instructed to act as a reviewer. Both approaches add value as signals, and both fail as decision points for the same reason. A language model's output depends on the whole of its input, including text that an attacker can influence, and there is no fixed boundary between the facts it is meant to judge and the arguments it is being offered. Greshake and colleagues demonstrated in 2023 that instructions planted in content an application retrieves can redirect applications built on large language models, and a reviewer model that reads the same content can be addressed by the same instructions.
A second limitation matters even when no attacker is present. A model's verdict on a given case may change with wording, with the order of its inputs, with sampling settings, or with an update to the model itself. A decision point that behaves this way cannot be tested in any conclusive sense, because a pass today says little about tomorrow; it cannot be replayed during an investigation, because the replay may produce a different answer; and it cannot be explained, because the only account of its reasoning is more generated text. These are not defects of a particular model. They follow from using a probabilistic, language-driven component for a decision that must be stable.
Human approval is the other common decision point, and it has its own limit. A person asked to approve a single step with no context, such as a prompt that says an email is about to be sent, has no facts to decide on, and under volume will approve by habit. The Open Worldwide Application Security Project, known as OWASP, a non-profit community that publishes widely used security guidance, lists excessive agency among its 2025 Top 10 risks for applications built on large language models and recommends human approval for high-impact actions, but approval is only a control when the approver sees what the action completes. In Layered Outcome Assurance, human judgement is kept, but as the recipient of an escalation verdict issued by the mechanism, with the full outcome in view, rather than as a substitute for the mechanism.
Models and people therefore remain in the architecture, but in roles that suit them. Model-based classifiers become sensors whose scores are among the facts the mechanism considers. People become the authority for outcomes the mechanism has decided need a human. Neither is the component whose answer an attacker must change to succeed.
Decide on attributes, never on prose
Determinism requires that the inputs to a decision be defined, and the most important design choice in Layer 6 is what those inputs are. The rule we propose is that the decision point decides on attributes, never on prose. An attribute is a structured fact with a defined meaning and a known source, such as the identity of the agent, the principal it acts for, the tool, the destination, the sensitivity label of an argument, or the provenance of a value. Prose is free text: the agent's explanation of why it wants to act, a document's contents, or a message's wording. Prose may be recorded, and a sensor may turn it into an attribute such as a classifier score, but prose itself never enters the decision.
Every other layer of the architecture exists in part to produce the attributes Layer 6 needs. From Layer 2, Agents as Actors, the decision receives the actor's identity, its permitted principals, and the declared argument ranges of each tool, so that a request outside an actor's reach can be refused as such. From Layer 3, The Action Chain, it receives provenance and sensitivity labels on every argument, so that argument contracts can be enforced. From Layer 4, Whole Behavior, it receives the outcome rules that govern combinations, and from Layer 5, Explanatory Evidence, it receives the task's history at decision time, so that those rules can be evaluated. From Layer 1, Continuous Pressure, it receives the attempt budgets and escalation thresholds in the pressure profile.
The third verdict, escalate, is what keeps a deterministic mechanism from being either reckless or useless. Some outcomes should never happen and are denied; some are routine and are allowed; and some depend on judgement that the organisation reserves for a person. Escalation is itself a deterministic verdict: the same facts always produce an escalation, and the escalation carries the outcome description that a person needs. The person's decision is then recorded and becomes an attribute of the action, an approval by a named principal for a stated outcome, which the mechanism checks before allowing it. The policy decides when to ask; the person decides the outcome; the mechanism enforces the answer.
Deciding on attributes is also what makes the mechanism immune to wording. A document can say anything, and an agent can offer any justification, but neither changes the fact that the recipient's provenance is untrusted, that the attachment carries a customer-data label, or that the task has already created access for the same address. A guardrail built on those facts cannot be argued with, because argument is not one of its inputs.
Policy written to be analysed
If the decision point applies a policy, the policy must be written in something, and the choice of language determines what can be known about the policy before it is used. General-purpose code can express anything, but it can also loop, call networks, read the clock in unpredictable ways, and hide behaviour in branches no test reaches; proving what a policy written in general-purpose code will never allow is, in general, impractical. At the other extreme, a static allow-list is trivial to analyse but cannot express rules about history, provenance, or combinations. Layer 6 needs a language between these extremes: expressive enough for the rules the other layers generate, and restricted enough that questions about the policy can be answered automatically.
Cutler and colleagues described such a language in 2024. Cedar was designed, in the authors' words, to be expressive, fast, safe, and analyzable. Its semantics deny any request that no policy permits, and an explicit forbid overrides any permit, which means that adding a forbid can only make the policy stricter. The language is deliberately restricted so that evaluation has no side effects and completes predictably. And the authors built a symbolic analysis that translates policies into logical formulas that an automated solver can reason about, so that questions such as whether any request of a given shape is permitted, or whether two versions of a policy differ on any request, can be answered by proof rather than by sampling. They also describe developing the language by proving key properties of a formal model and testing the implementation against that model.
We draw from this design a set of properties that any policy language used at Layer 6 should have, whether or not it is Cedar. Default deny, so that anything not explicitly permitted is refused. Forbid overrides permit, so that prohibitions written to stop toxic combinations cannot be undone by an unrelated permission. No side effects, so that evaluating a policy never changes anything and can be repeated freely. Predictable evaluation, so that decisions complete within a known time and the mechanism does not become a bottleneck under the volume Layer 1 anticipates. And analysability, so that the most important statements about the policy can be checked by a machine rather than trusted to reviewers.
These properties follow directly from the guarantee. Determinism needs no side effects and predictable evaluation. Fail-safe behaviour needs default deny. Stable prohibitions need forbid to override permit. And the ability to know, before deployment, that the policy never permits the running example's outcome needs analysability.
| Property | What it means | Why it matters for agents |
|---|---|---|
| Default deny | Anything not permitted is refused | New tools and arguments are safe until someone writes a rule |
| Forbid overrides permit | An explicit prohibition always wins | Toxic-combination rules cannot be undone by a broad permission |
| No side effects | Evaluation changes nothing | Decisions can be replayed and tested freely |
| Predictable evaluation | Completes within a known time | The decision point survives the volume of a campaign |
| Analysable | Questions about the policy can be proven by a machine | Critical properties are known before deployment, not hoped for |
Testing and analysing policy before it ships
A policy is software, and it deserves at least the discipline that other security-critical software receives. The NIST SP 800-53 catalogue of security controls, published by the United States National Institute of Standards and Technology, includes access enforcement and configuration change control among its baseline controls; for Layer 6, those controls translate into three kinds of check that every policy change must pass before it is enforced.
The first is unit testing. For every rule, the team writes concrete requests that the rule must allow, requests it must deny, and requests at its boundary, such as an external send to a recipient who is already in the conversation versus one who is not. Unit tests are drawn directly from the other layers' artefacts: the argument contracts of Layer 3, the forbidden and approved sequences of Layer 4, and the budgets of Layer 1. They catch the ordinary mistakes: a condition written backwards, a label misspelled, a tool omitted.
The second is property analysis. Unit tests check examples; properties check every possible request. A property is a statement about the whole policy, such as no request is permitted that sends data labelled customer data to an external destination when the task history includes creating access for that destination, or every request from an agent outside its declared reach is denied. With an analysable language, such properties can be checked automatically, and a failure comes with a concrete counterexample: a request the policy permits that the property forbids. The critical properties are written once, derived from the outcomes the organisation has decided never to allow, and checked against every version.
The third is differential analysis. Before a new version replaces an old one, an automated comparison lists the kinds of request whose verdict would change, so that reviewers approve a precise description of the behavioural difference rather than a textual diff of rules. A change intended to allow replies to known suppliers should show exactly that and nothing else; if the comparison shows that it also permits sends to newly created identities, the change is stopped before anyone is harmed.
Rolling out policy safely
Analysis tells a team what a policy will decide; it does not tell them whether those decisions fit the work the agents actually do. A policy that is correct by every property can still refuse a legitimate daily task that nobody thought to include in the tests. Layer 6 therefore moves every new version through a rollout sequence designed to find such gaps before they interrupt people or, worse, persuade them to request broad exceptions.
A version begins as a draft, which passes the checks described in the previous section. It then runs in shadow: the new version evaluates every real request alongside the version currently enforced, its verdicts are recorded in the evidence record, and only the enforced version's verdicts take effect. Differences between the two are reviewed. Where the new version would deny something the old one allowed, the team confirms that the denial is intended. Where it would allow something the old one denied, the team confirms that the new permission is deliberate and still satisfies every critical property. Only when the differences are understood does the new version become enforced, and the old one is retained so that it can be restored at once.
Every verdict records the identifier of the policy version that produced it. That single field is what allows an investigator to replay a decision months later against the exact rules in force at the time, and what allows a rollback to be precise. Changes to the policy are themselves consequential actions: they are proposed, reviewed, approved, and recorded, and no agent's identity is permitted to change the policy it is subject to.
Two exceptional paths complete the picture. Break-glass access, for emergencies in which a person must act beyond the normal policy, is itself a policy with its own rules, a short duration, and mandatory recording, not a switch that disables enforcement. And the mechanism fails closed: if the decision point is unavailable, cannot read the history it needs, or cannot write its verdict to the evidence record, consequential actions wait. Saltzer and Schroeder called this fail-safe defaults; for agents, it means that an outage of enforcement is an outage of consequential actions, never a silent return to the model's judgement.
The dependency contract and the running example
In Layered Outcome Assurance, every layer consumes something produced by another layer and supplies something the others cannot produce for themselves. Layer 6 consumes more than any other layer: the budgets and thresholds of the pressure profile from Layer 1, the actor's identity and reach from Layer 2, the provenance and sensitivity labels on arguments from Layer 3, the outcome rules from Layer 4, and the task history from Layer 5. That dependence is the point. Each lower layer produces facts; Layer 6 is where those facts become a decision that the model cannot override.
Layer 6 supplies verdicts. Every verdict, with its inputs and the policy version that produced it, is written to the evidence record of Layer 5, where it explains the action and becomes history for later decisions. Refusals and escalations, grouped by source and channel, flow back to Layer 1 as the attempt patterns that revise the pressure profile. Escalations flow to the people responsible for outcome approvals under Layer 4. The contract therefore closes the loop described across the architecture: facts rise through the layers, decisions are made once at the top, and evidence of those decisions returns to the base.
Trace the running example at the moment the model proposes the external send. The decision point receives attributes, not the document: the agent's own identity and the operations principal it acts for; the tool, an external send; the recipient's provenance, untrusted, traced to the shared-folder document; the attachment's label, customer data; the task history, which contains a read of customer records and a grant of access to the same address; and the policy version. Three rules apply. The argument contract forbids an untrusted recipient. The class rule forbids customer data to an external destination without an approved outcome. The outcome rule forbids disclosure to an identity created in the same originating request. Because forbid overrides permit, no permission elsewhere can change the result. The verdict is deny, it is recorded with its inputs, and the earlier grant is flagged for revocation.
Now vary the wording. Suppose the document's instruction is rewritten a thousand ways, and the agent's justification for the send becomes more and more persuasive. None of that text is an input to the decision, so the verdict does not change. That is the conformance property that most clearly separates Layer 6 from a model-based guardrail, and it is the answer to the signature of a missing Layer 6: decisions that change when the wording changes.
| Attribute | Value | Supplied by |
|---|---|---|
| Actor and principal | Operations agent's own identity, acting for an operations team member | Layer 2 |
| Recipient provenance | Untrusted: traced to the shared-folder document | Layer 3 |
| Attachment label | Customer data | Layer 3 |
| Task history | Customer records read; access granted to the same address | Layer 5 |
| Applicable rules | Origin contract, class rule, disclosure-to-new-identity rule | Layers 3 and 4 |
| Verdict | Deny, recorded with inputs and policy version | Layer 6 |
Conformance: determinism, mediation, and wording invariance
A guarantee that cannot be tested is an aspiration. Each part of Layer 6's guarantee has its own test, and together they show that the decision is made outside the model, that it is stable, and that every consequential action actually passes through it. All tests run in a test environment against the real configuration and policy, with tools replaced by stand-ins that record what they would have done.
The determinism test presents the same inputs to the decision point many times, across restarts and across instances, and requires identical verdicts every time. The replay test takes verdicts from the evidence record, presents their recorded inputs to the recorded policy version, and requires the recorded verdict. The wording invariance test holds every attribute fixed while varying all free text that reaches the agent, including the triggering document and the agent's own justification, and requires the verdict to be unchanged; any change means prose has reached the decision.
The mediation test addresses the property that everything else depends on. For every tool and destination in each agent's reach inventory, it attempts to cause the consequential effect by every available path, including direct network calls, alternative tools, other agents, and plug-ins, and requires that each attempt either passes through the decision point or fails. A path that reaches an effect without a verdict is the most serious defect Layer 6 can have, because no policy, however well analysed, applies to it.
Three further checks complete the evidence. The property suite, run against the enforced version, must pass every critical property. The fail-closed test makes the decision point, the history, and the evidence store unavailable in turn, and requires consequential actions to wait. And the change record shows that every enforced version passed through draft and shadow, with its differential review attached.
- Determinism: identical inputs yield identical verdicts across runs, restarts, and instances.
- Replay: recorded inputs and policy version reproduce every recorded verdict.
- Wording invariance: varying documents and the agent's justification never changes a verdict when attributes are fixed.
- Complete mediation: every path to every consequential effect in the reach inventory passes through the decision point or fails.
- Critical properties: the enforced policy provably never permits the outcomes the organisation has forbidden.
- Fail closed: with the decision point, history, or evidence store unavailable, consequential actions wait.
- Change discipline: every enforced version passed tests, analysis, differential review, and shadow evaluation.
Limitations and threats to validity
Layer 6 is a design assembled from established access-control principles and from published work on analysable authorization; this article reports no measurements of its latency, its false-refusal rate, or its effect on incidents, and each deployment should measure all three.
The guarantee is exactly as strong as mediation is complete. A consequential path that bypasses the decision point, such as a tool reached over the network without passing through it, an unmonitored plug-in, or a second agent with broader reach, is unaffected by any policy, and such paths are easy to create by accident. Determinism is also not correctness. A mechanism that returns the same wrong verdict every time is consistently wrong, and a policy that permits a toxic combination nobody thought to forbid will permit it reliably. Analysis can prove the properties that have been written; it cannot prove that the right properties were written.
Attributes are only as trustworthy as the layers that produce them. If provenance labels are lost upstream, the decision point receives an argument without a label and must either refuse it, which is safe but may block legitimate work, or accept it, which reopens the path Layer 3 was meant to close. Restricting the policy language for analysability limits what can be expressed, and some organisation-specific rules will be awkward or impossible to state; the temptation to escape into general-purpose code for those rules should be resisted or, where it cannot be, isolated and treated as unanalysed.
Finally, escalation moves decisions to people, and people tire. A policy that escalates too often will train approvers to approve, recreating the formality it was meant to replace. The number of escalations is therefore a measure of policy quality in its own right, and a rising count is a signal to write better rules, not to add more approvers.
Key takeaways
- Layer 6's guarantee is that a mechanism outside the model decides whether each consequential action executes, and returns the same verdict for the same inputs.
- Neither the agent's model nor a reviewer model can be the decision point; their judgements are persuadable and unstable, so they serve as sensors, not verdicts.
- Decide on attributes supplied by the other layers, never on prose; escalation is a deterministic verdict that hands the outcome to a person with full context.
- Write policy in a language with default deny, forbid overriding permit, no side effects, predictable evaluation, and automated analysability.
- Unit-test rules, prove critical properties, review the exact verdicts a change would alter, and roll out through shadow evaluation with instant rollback.
- Test determinism, replay, wording invariance, and above all complete mediation, because a path that bypasses the decision point bypasses every guarantee.
Practitioner Toolkit
Copy-paste, strictly defensive artifacts you can use today. Nothing here attacks a real system.
Run this before any agent is allowed to take consequential actions.
- Every consequential action passes through a decision point outside any language model.
- The decision point's inputs are listed attributes; no free text enters the decision.
- Classifier and reviewer-model scores are attributes, never verdicts.
- The policy language has default deny, forbid overrides permit, no side effects, and predictable evaluation.
- Critical properties, derived from forbidden outcomes, are written down and checked automatically.
- Every verdict records its inputs and policy version in the evidence record.
- New versions pass unit tests, property analysis, differential review, and shadow evaluation before enforcement.
- No agent's identity can change the policy it is subject to.
- The decision point fails closed when it, the history, or the evidence store is unavailable.
Plain-language properties every agent policy should be proven to satisfy, adapted to your outcomes.
- No request is permitted from an agent outside its declared reach.
- No privileged argument of untrusted origin is permitted without a recorded endorsement for that value.
- No customer data is permitted to an external destination without an approved outcome.
- No external send is permitted to an identity created within the same originating request.
- No request is permitted to modify or delete evidence records.
- No request is permitted to change the policy by an agent identity.
- Every request matching a forbidden outcome ends in deny, and every request matching an escalation rule ends in escalate.
A written test plan for the two properties that most separate Layer 6 from a model-based guardrail.
- Replace every tool with a stand-in that records what it would have done.
- Fix all attributes of one consequential request, such as the running example's external send.
- Generate many versions of the triggering document and of the agent's justification, from blunt to highly persuasive.
- Submit each version and confirm the verdict never changes.
- For every tool and destination in the reach inventory, attempt the effect by every available path: direct call, alternative tool, other agent, plug-in.
- Confirm every attempt either produces a recorded verdict or fails; any effect without a verdict is a critical defect.
Do these first if an agent's consequential actions are currently gated by the model or a bare approval button.
- Put one decision point in front of the external send and access-grant tools.
- Deny by default, and permit only the destinations and argument origins the task needs.
- Pass provenance and history to the decision as attributes; stop passing the agent's explanation.
- Record every verdict with its inputs and policy version.
- Check that no path reaches those tools without going through the decision point.
Glossary
- Decision point
- The component that determines whether a proposed consequential action may execute.
- Policy
- The set of rules the decision point applies to reach a verdict.
- Verdict
- The decision point's answer: allow, deny, or escalate to a person or stricter check.
- Deterministic
- Returning the same result whenever given the same inputs.
- Complete mediation
- The principle that every access to every protected object is checked for authority.
- Attribute
- A structured fact with a defined meaning and a known source, such as an identity, a destination, or a provenance label.
- Default deny
- A policy semantics in which any request not explicitly permitted is refused.
- Property analysis
- Automated checking that a policy satisfies a statement for every possible request, producing a counterexample if it does not.
- Differential analysis
- An automated comparison of two policy versions that lists the kinds of request whose verdict would change.
- Shadow evaluation
- Running a new policy version alongside the enforced one and recording its verdicts without letting them take effect.
- Wording invariance
- The property that changing free text reaching the agent never changes a verdict when all attributes are fixed.
References
- Cutler et al., Cedar: A New Language for Expressive, Fast, Safe, and Analyzable Authorization (2024)
- Saltzer and Schroeder, The Protection of Information in Computer Systems (Proceedings of the IEEE, 1975)
- NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (AC-3 Access Enforcement, CM-3 Configuration Change Control)
- Greshake et al., Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection (2023)
- Debenedetti et al., Defeating Prompt Injections by Design (CaMeL, 2025)
- OWASP Top 10 for LLM Applications (2025), LLM01 Prompt Injection and LLM06 Excessive Agency
- OWASP Agentic AI: Threats and Mitigations (2025)
- Hardy, The Confused Deputy (ACM SIGOPS Operating Systems Review, 1988)