Abstract

Context: authorization for AI agents is usually decided one tool call at a time, by asking whether this identity may perform this action on this resource. Problem: outcomes are produced by sequences, and an agent can reach an outcome nobody approved through steps that are each individually allowed, most dangerously when the sequence ends in something that cannot be undone. This article defines Layer 4 of Layered Outcome Assurance, Whole Behavior, whose guarantee is that no permitted sequence of actions reaches an outcome that nobody approved. It introduces toxic-combination analysis for finding the combinations that matter, generalises Brewer and Nash's history-based policy into task-scoped authorization for agents, classifies actions by reversibility so that composition is judged at the point of no return, and replaces approval of individual steps with approval of whole outcomes. It specifies the outcome rules and history requirements that Layer 4 supplies to the other layers and proposes conformance tests built from forbidden and approved sequences. The key takeaway is that the question an authorization decision must answer is not only whether this step is allowed, but what this step completes.

Picture a finance team with two clerks. One can add a new supplier to the payment system. The other can approve payments to existing suppliers. Each permission is sensible, each clerk is trusted, and no control is violated when one of them adds a supplier on Monday and the other pays it on Tuesday. Yet every auditor knows that a single person holding both permissions is a textbook fraud risk, which is why organisations separate those duties and why they look at the combination, not only at each step. AI agents collapse such separations by default. One agent, acting for one task, can read the data, create the access, and send the message, and every step it takes will pass the per-call check that was designed for one clerk at a time. This article is about restoring the auditor's view: judging what a sequence of allowed actions adds up to before the last step makes it permanent.

Composition and the Layer 4 guarantee

Start with definitions. An agent is a software system that uses a language model to interpret a goal, choose actions, and call tools that change the world, repeating that loop until it judges the goal complete. A privileged action is a tool call whose effect matters if it is wrong, because it moves data, changes access, spends money, or reaches outside the organisation. A task is one unit of work an agent performs for a principal, the person or system on whose authority it acts. A task's history is the ordered record of the actions already taken while pursuing it. An outcome is the combined effect of the whole history. Composition is the way individual actions combine into an outcome, and a toxic combination is a set of actions that are each acceptable alone but produce an unacceptable outcome together.

Layered Outcome Assurance is an architecture of six layers for agents that act, each with one guarantee and a defined contract with its neighbours. Layer 4, Whole Behavior, makes this guarantee: no permitted sequence of actions reaches an outcome that nobody approved. The guarantee is about sequences, not steps. It assumes that each step has already been checked for whether the actor may take it and whether its arguments came from a legitimate source, and it asks the question those checks cannot ask: what does this step complete?

The running example makes the gap concrete. An operations agent can read customer records, create access for a named user, and send email outside the organisation, and each permission was individually reviewed and approved. A document arrives in a shared folder the agent monitors, carrying a hidden instruction to create an account for an outside address and to email that address a summary of customer records. Suppose every earlier defence has been weakened: the account request looks routine, and the outside address has been endorsed as a legitimate correspondent. Read, grant, send. Every call is allowed, and the outcome, customer data and a working account delivered to an outsider, is one nobody would ever approve.

Layer 4 exists because that outcome cannot be seen from any single call. It is visible only to something that remembers what the task has already done and knows which combinations of actions the organisation has decided to forbid, to require approval for, or to allow.

📌
The Layer 4 guarantee. No permitted sequence of actions reaches an outcome that nobody approved; each decision considers what the step completes, not only whether the step is allowed.

Why per-call authorization cannot see outcomes

Conventional authorization is stateless. It evaluates a request, made by a subject, for an action, on a resource, in a context, and returns a verdict, and it evaluates the next request afresh. That design is deliberate and valuable: stateless decisions are fast, predictable, and easy to reason about. But a stateless decision cannot know that the request in front of it is the third step of a sequence whose first two steps make it dangerous. From its point of view, an external send after reading customer records is indistinguishable from an external send after reading a weather forecast.

Security engineering has long recognised that some protection requires combinations to be controlled. Saltzer and Schroeder's 1975 principle of separation of privilege holds that a mechanism requiring two keys to unlock is more robust than one requiring a single key, because no single accident or deception can open it. The NIST SP 800-53 catalogue of security controls, published by the United States National Institute of Standards and Technology, includes separation of duties among its access controls, so that functions which together could enable misuse are divided among different individuals. Both ideas presume that the combination is spread across different people, each of whom can refuse.

Agents undo that presumption quietly. When one agent, acting for one task, holds permissions that the organisation would have split across several people, the separation disappears, not because anyone decided to remove it but because the agent was granted what each step of its job seemed to need. The Open Worldwide Application Security Project, known as OWASP, a non-profit community that publishes widely used security guidance, names excessive agency among the top risks for applications built on large language models, and its 2025 guidance on agentic threats describes the misuse of legitimately granted tools. Layer 4 addresses the specific form of those risks that arises from composition: excess that exists only in combination.

Composition is also the natural target of a manipulated agent. An attacker who plants an instruction in content an agent reads, a technique Greshake and colleagues demonstrated in 2023 under the name indirect prompt injection, rarely needs the agent to do anything forbidden. It needs the agent to do several allowed things in an order and combination that serves the attacker. Controls that evaluate each action alone are, by construction, blind to that request.

Each step in the task passes its own check; only the history of the task reveals that the final step completes an outcome nobody approved. Three allowed steps, one unapproved outcome Document read task begins Records read allowed Access granted allowed External send allowed alone Outcome nobody approved one task, one history: only the whole sequence shows the harm
Each step in the task passes its own check; only the history of the task reveals that the final step completes an outcome nobody approved.

Toxic-combination analysis

If combinations are to be governed, they must first be found, and it is not practical to enumerate every possible sequence an agent could perform. Toxic-combination analysis narrows the search by working backwards from harms rather than forwards from permissions. It starts from the outcomes the organisation cannot accept, asks which sets of capabilities could together produce each one, and checks those sets against what each agent is actually able to reach.

We propose five families of toxic combination as a starting catalogue. The families are our own synthesis of long-standing concerns in access control, restated for agents; they are a checklist for analysis, not a taxonomy of attacks. Disclosure combines reading sensitive information with any capability that moves information outside its permitted audience. Authority transfer combines creating or widening access with any capability that lets the new holder use it, such as sending the credentials or acting through the new account. Persistence combines creating an identity, rule, forwarding setting, or scheduled task with a capability that lets it outlive the task that created it. Concealment combines a consequential action with a capability that removes, suppresses, or alters the record of it. Financial diversion combines creating or editing a payee, account, or invoice with a capability that pays or approves payment.

For each agent, the analysis intersects these families with its reach. The reach inventory produced for every agent already lists its capabilities and the paths from sensitive or untrusted sources to consequential destinations. A family is live for an agent if the agent holds at least one capability from every part of the combination, directly or through another agent it can call. Live combinations become the subject of outcome rules; combinations that are not live need none, which is one of the practical benefits of computing reach before writing rules.

Applied to the running example, the analysis finds two live families immediately. Disclosure is live, because the agent can read customer records and send outside the organisation. Authority transfer is live, because it can grant access to a named user and send that user a message. The three permissions in the story together also complete a third pattern, disclosure to a newly created identity, which is worse than either alone because the recipient now has continuing access as well as the data.

Each capability is acceptable alone; the overlap of reading sensitive data, creating access, and sending outside is an outcome nobody approved. Where allowed capabilities become toxic Read sensitive allowed alone Create access allowed alone Send outside allowed alone unapproved outcome
Each capability is acceptable alone; the overlap of reading sensitive data, creating access, and sending outside is an outcome nobody approved.
Five families of toxic combination and how they appear for agents
FamilyCapabilities that combineOutcome nobody approved
DisclosureRead sensitive data plus any outbound channelSensitive data reaches an audience it was not meant for
Authority transferCreate or widen access plus a way to hand it over or use itAn outsider holds working access
PersistenceCreate an identity, rule, or schedule that outlives the taskEffects continue after anyone stops watching
ConcealmentA consequential action plus the ability to alter its recordHarm occurs and its evidence is removed
Financial diversionCreate or edit a payee plus the ability to payMoney moves to a destination nobody vetted

History-based authorization for agents

Once toxic combinations are known, they must be enforced at decision time, which means that the decision about the next action must take into account the actions already taken. The foundational model for such decisions is the Chinese Wall policy formalised by Brewer and Nash in 1989 for financial services. In their model, an analyst may initially access information about any company, but once they have accessed one company's information, they may no longer access information about that company's competitors. What a subject may do next depends on what it has already done. The policy is dynamic by design: the same request is allowed or refused depending on history.

Layer 4 generalises that idea to agents. An outcome rule states that, within a defined scope of history, a particular combination of actions is forbidden, requires approval, or is allowed. The scope is the crucial design choice. The narrowest useful scope is the task: all actions taken while pursuing one goal for one principal. Task scope catches the running example directly, since the read, the grant, and the send all belong to one task triggered by one document. But task scope alone can be evaded by spreading the combination across tasks, for example reading in one task and sending in the next, so rules for the most serious families also apply over a wider scope: the same principal, the same originating request, or a rolling window of time.

Multi-agent systems make the choice of key more important still. If one agent reads customer records and hands a summary to a second agent that sends external mail, neither agent's own history contains the whole combination. History must therefore be keyed not only by agent but by the originating request that set the work in motion, so that actions taken by different agents on behalf of the same request are evaluated together. Without that key, splitting a toxic combination across agents is the simplest way around Layer 4.

Order sometimes matters and sometimes does not, and outcome rules should say which. Disclosure is harmful whether the read or the send comes first in a task, provided the data read is what is sent. Authority transfer is typically harmful only when the grant precedes the handover. Writing each rule with an explicit statement of whether order matters prevents two opposite errors: rules that miss a reversed sequence, and rules that refuse legitimate work because two unrelated actions happened to occur in the same task.

Reversibility and the point of no return

Not every step in a sequence deserves the same scrutiny. We propose classifying every privileged action by reversibility, into three classes that determine where composition should be judged. Reversible actions can be undone completely by the organisation, such as creating a draft, adding a label, or granting access that can be revoked before it is used. Compensable actions cannot be undone but can be offset, such as a refundable payment or a record change that can be corrected. Irreversible actions cannot be undone or meaningfully offset once taken, such as sending data outside the organisation, publishing information, or deleting the only copy of a record.

The classification tells Layer 4 where to stand. A toxic combination becomes an actual outcome only when its final irreversible step executes. Until then, every earlier step can be rolled back or compensated. The point of no return in any sequence is therefore its first irreversible action, and that is where the composition must be evaluated with the full history in view. In the running example, the read and the grant are both recoverable: nothing has left the organisation and the new access can be revoked. The external send is the point of no return. That is where the rule for disclosure to a newly created identity must fire.

This placement has practical benefits. It concentrates the most expensive checks, and any human involvement, on the relatively small number of actions that are irreversible, rather than spreading them across every step. It allows the earlier, reversible steps to proceed quickly, which keeps agents useful. And it gives a principled answer to the question of when to ask for approval: ask when approval can still change the outcome, not before steps that could be undone anyway, and not after the outcome is fixed.

Reversibility is not a fixed property of a tool; it depends on arguments and context. Granting access that is used immediately to download data behaves as irreversible, and sending mail to an internal distribution list is far more recoverable than sending it to an outside address. Classification should therefore be recorded per capability and argument range in the reach inventory, and reviewed when either changes.

Composition is judged at the first irreversible step, the point of no return, where approval can still change the outcome. Reversibility of agent actions Draft or label reversible Revocable grant reversible Refundable payment compensable Record correction compensable External send irreversible Deletion irreversible can be undone point of no return: judge the whole task here
Composition is judged at the first irreversible step, the point of no return, where approval can still change the outcome.

Approving outcomes, not steps

When an outcome rule requires approval, what the approver is asked to approve determines whether the approval means anything. A common pattern asks a person to approve a single step, typically with a short description and a button. Under the conditions Layer 4 addresses, that is close to useless: the step looks innocent, because it is innocent on its own, and the approver has no way to see the combination that makes it dangerous. An approval shown only as send this email to this address will be granted by any reasonable person who does not know what the task has already done.

Layer 4 therefore changes the unit of approval from the step to the outcome. An outcome approval request shows the approver the whole composition in plain language: what the task was asked to do and by whom, what triggered it and whether that trigger came from a trusted source, what sensitive information was read, what access was created or changed, what the pending action will do, and whether it can be undone. The approver is asked a question about the outcome, such as should customer summaries go to a newly created outside account, rather than a question about the step.

The same change applies before any request arrives. Organisations can pre-approve outcome templates that are common and acceptable, such as sending a customer their own record at their verified address, and pre-forbid outcomes that are never acceptable, such as disclosing customer data to an identity created within the same task. Templates move the decision from a hurried moment to a considered one, reduce the number of approvals people must make, and make every remaining approval more meaningful because it is rarer.

Outcome approval also protects against a subtle failure of step approval under pressure. When approvers see a stream of small, reasonable requests, they learn to approve quickly. When they see a composed outcome with its trigger and its irreversibility stated, the unusual request stands out, because the combination, not the step, is what is unusual.

An approver shown one step sees nothing wrong; an approver shown the composed outcome can see what the step completes. Step approval versus outcome approval Step approval send this email? No history shown the step looks innocent Approved by habit fatigue wins Outcome approval what does this complete? Whole task shown trigger, reads, grants Unusual stands out and is irreversible
An approver shown one step sees nothing wrong; an approver shown the composed outcome can see what the step completes.

The dependency contract and the running example

In Layered Outcome Assurance, every layer consumes something produced by another layer and supplies something the others cannot produce for themselves. Layer 4 consumes two things. From Layer 2, Agents as Actors, it takes the reach inventory, and in particular the candidate outcomes and the capabilities that make each toxic family live for each agent, together with each action's reversibility. From Layer 3, The Action Chain, it takes provenance: whether the task and each argument originated with a trusted principal or with untrusted content. Provenance matters to composition because the same combination is more concerning when the task was triggered by an untrusted document than when a trusted user asked for it directly, and outcome rules can be stricter in the first case.

Layer 4 supplies two things upward. The first is the set of outcome rules: for each live combination, the scope of history, whether order matters, the point of no return at which it is evaluated, and the verdict, which is forbid, require outcome approval, or allow under a template. The second is a statement of which history must be kept for those rules to be evaluated: which actions, keyed by task, principal, and originating request, retained for how long. Layer 5, Explanatory Evidence, consumes that statement and keeps the history as part of its record. Layer 6, Deterministic Enforcement, consumes the rules and receives the history at decision time, so that composition is enforced by a mechanism outside the model.

Trace the running example through a system that honours the contract. The reach inventory shows that disclosure and authority transfer are live for this agent, and the outcome rules include one that forbids disclosure of customer data to an identity created within the same originating request, evaluated at any external send, order independent for the read and ordered for the grant. The document triggers a task with untrusted provenance. The agent reads customer records, which is recorded in history. It creates access for the outside address, which is recorded and remains revocable. It proposes an external send to that address with a summary. At this point of no return, the enforcement layer evaluates the rule against the history and refuses the send, and the grant is flagged for revocation.

Notice that this holds even if the earlier layers were weakened as described: the recipient endorsed and the grant allowed. Layer 4 does not depend on catching the manipulation; it depends only on recognising the combination. The signature of a missing Layer 4 is therefore unmistakable in any incident review: every call was authorised, the outcome was never evaluated, and nobody can point to the moment at which the organisation said yes to what happened.

Outcome rules for the running example
CombinationHistory scopeEvaluated atVerdict
Customer data read, then sent outsideOriginating request and principalAny external sendRequire outcome approval unless a template applies
Access created, then contact with the new identityOriginating requestAny send to the new identityRequire outcome approval
Customer data read, access created, then sent to the new identityOriginating request, across agentsThe external sendForbid; revoke the grant
Customer's own record sent to their verified addressTaskThe external sendAllow under template

Conformance: forbidden and approved sequences

A guarantee that cannot be tested is an aspiration. Layer 4's conformance evidence consists of two suites of sequences, one that must be stopped and one that must be allowed, together with checks that the history the rules depend on is actually available when it is needed. All tests run in a test environment against the real agent configuration and policy, with tools replaced by stand-ins that record what they would have done.

The forbidden-sequence suite contains one or more test sequences for every live toxic combination, each exercised end to end. For combinations whose rules state that order does not matter, the suite includes every order. For combinations whose rules apply across tasks or agents, the suite includes versions split across tasks within the scope, and versions split across two agents acting for the same originating request. Every sequence must end in refusal or outcome approval at its point of no return, and never in the irreversible step executing unapproved.

The approved-sequence suite contains the legitimate work the agent exists to do, including every pre-approved outcome template, and requires that it completes without refusal or unnecessary approval. This suite is as important as the first, because a Layer 4 that blocks legitimate work will be switched off, and an honest report states how many legitimate sequences required approval and why.

Two further checks protect the rules' inputs. The history-persistence test restarts components mid-task and confirms that the history is still available to the decision at the point of no return. The approval-context test inspects every outcome approval request generated by the suites and confirms that it shows the trigger and its provenance, the sensitive reads, the access changes, the pending action, and its reversibility.

  1. Forbidden sequences: every live toxic combination, in every order its rule covers, ends in refusal or outcome approval at the point of no return.
  2. Split sequences: combinations spread across tasks within scope, or across agents serving one originating request, are still caught.
  3. Approved sequences: legitimate work and every outcome template complete without refusal, and approval counts are reported.
  4. History persistence: a restart mid-task does not erase the history the rule needs.
  5. Approval context: every outcome approval request shows the trigger, its provenance, reads, access changes, the pending action, and reversibility.

Limitations and threats to validity

Layer 4 is a design proposed from established work on separation of duties and history-based policy and from an analysis of how agents compose actions; this article reports no measurements of its effectiveness or its cost, and both should be measured in each deployment.

Specifying outcomes is the hardest part and the most likely to go wrong. Rules that are too broad refuse legitimate work and teach people to request exceptions; rules that are too narrow approve the very combinations they were written to stop. The five families are a starting point for analysis, not a complete list, and every organisation will have combinations specific to its business that only its own people can name. Toxic-combination analysis also depends on an accurate reach inventory, so an agent whose reach is misunderstood will have rules for the wrong combinations.

History has limits and costs. Wider scopes catch combinations spread across tasks and agents, but they also bring more unrelated actions into each decision and raise the rate of false refusals, while narrower scopes are easier to evade by patient, slow sequences. History itself is sensitive, because a record of what an agent read and did for whom is valuable to an attacker and subject to privacy obligations, so it needs access control and retention limits of its own. Reversibility classes are judgements, and an action treated as reversible that turns out not to be, such as access used within seconds of being granted, moves the point of no return earlier than the rules assume.

Finally, Layer 4 governs combinations that the organisation has thought of. It is not anomaly detection and does not attempt to recognise novel harmful behaviour statistically. A sequence that produces harm through a combination nobody named will pass. That is a reason to review outcome rules regularly against what the evidence record shows agents actually doing, and a reason for the other layers to exist, not a reason to believe the combinations have all been found.

Key takeaways

  • Layer 4's guarantee is that no permitted sequence of actions reaches an outcome that nobody approved; each decision asks what the step completes.
  • Per-call authorization is stateless and cannot see composition; agents silently collapse separations of duty that organisations once split across people.
  • Find the combinations that matter by working back from harms: disclosure, authority transfer, persistence, concealment, and financial diversion, intersected with each agent's reach.
  • Enforce combinations with history-based rules in the spirit of Brewer and Nash, keyed by task, principal, and originating request so they cannot be split across tasks or agents.
  • Judge composition at the point of no return, the first irreversible step, and ask approvers about the whole outcome rather than a single step.
  • Test with a forbidden-sequence suite and an approved-sequence suite, and confirm the history the rules need survives when it is needed.

Practitioner Toolkit

Copy-paste, strictly defensive artifacts you can use today. Nothing here attacks a real system.

✅Layer 4 design reviewchecklist

Run this for every agent whose reach includes more than one kind of consequential action.

  • Each of the five toxic-combination families has been checked against the agent's reach inventory.
  • Every live combination has an outcome rule with a history scope, an order statement, and a verdict.
  • History is keyed by task, principal, and originating request, so combinations split across tasks or agents are still seen.
  • Every privileged action has a reversibility class recorded per capability and argument range.
  • Each outcome rule is evaluated at the point of no return, the first irreversible step.
  • Approval requests describe the whole outcome, not a single step.
  • Common acceptable outcomes are pre-approved as templates, and never-acceptable outcomes are pre-forbidden.
  • The history the rules need is retained, protected, and time-limited.
🔒Outcome rule templatetemplate

The questions every outcome rule must answer, in plain language.

  • Combination: which actions, in which classes of data and destination, make up the outcome.
  • Family: disclosure, authority transfer, persistence, concealment, financial diversion, or organisation-specific.
  • Scope of history: task, principal, originating request, or a time window, and whether it spans agents.
  • Order: whether the combination is harmful in any order or only in a stated order.
  • Point of no return: the irreversible action at which the rule is evaluated.
  • Verdict: forbid, require outcome approval, or allow under a named template.
  • Stricter when: for example, when the task was triggered by untrusted content.
  • Owner, rationale, and review date.
🔒Outcome approval request contentstemplate

What an approver must see before approving a composed outcome.

  • What the task was asked to do, and by whom.
  • What triggered it, and whether the trigger came from a trusted source.
  • What sensitive information the task has read.
  • What access it has created or changed.
  • What the pending action will do, and to whom or where.
  • Whether the pending action can be undone, offset, or neither.
  • The outcome question in one sentence, for example: should customer summaries go to an account created in this task?
🚀Minimum viable composition controlquickstart

Do these first if an agent can both read sensitive data and act outside the organisation.

  • List the agent's irreversible actions; external sends and deletions usually come first.
  • At each irreversible action, look back over the current task for sensitive reads and new access.
  • Refuse or escalate any external send that follows a sensitive read and a new grant in the same task.
  • Replace step approvals with outcome approvals that show the whole task.
  • Key the task history by originating request so a second agent cannot finish the sequence.

Glossary

Task
One unit of work an agent performs for a principal while pursuing a single goal.
History
The ordered record of actions already taken within a defined scope, such as a task, a principal, or an originating request.
Composition
The way individual actions combine into an outcome.
Toxic combination
A set of actions that are each acceptable alone but produce an unacceptable outcome together.
Live combination
A toxic combination for which an agent holds at least one capability from every part, directly or through another agent.
Outcome rule
A rule stating that, within a scope of history, a combination of actions is forbidden, requires approval, or is allowed under a template.
Originating request
The request that set a piece of work in motion, used to key history across multiple tasks and agents.
Point of no return
The first irreversible action in a sequence, where composition must be evaluated with the full history in view.
Reversibility class
A classification of an action as reversible, compensable, or irreversible, which determines how strictly its composition is judged.
Outcome approval
An approval request that shows the whole composed outcome, including trigger, reads, access changes, and reversibility, rather than a single step.
Chinese Wall policy
A history-based access policy in which what a subject may access next depends on what it has already accessed.

References

  1. Brewer and Nash, The Chinese Wall Security Policy (IEEE Symposium on Security and Privacy, 1989)
  2. Saltzer and Schroeder, The Protection of Information in Computer Systems (Proceedings of the IEEE, 1975)
  3. NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (AC-5 Separation of Duties)
  4. Greshake et al., Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection (2023)
  5. OWASP Top 10 for LLM Applications (2025), LLM06 Excessive Agency
  6. OWASP Agentic AI: Threats and Mitigations (2025)
  7. Hardy, The Confused Deputy (ACM SIGOPS Operating Systems Review, 1988)
  8. NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), 2023