Research seriesL3paper
Securing an AI agent one component at a time is no longer enough: every component can pass review while the agent's overall behavior reaches an outcome nobody approved. This series defines Layered Outcome Assurance, an architecture of six layers for agents that act. Each layer makes one guarantee true, depends on the layer below it, and supplies something the layer above cannot work without: design for continuous pressure, bound every agent's reach as an actor, keep untrusted content away from privileged actions, govern whole sequences of behavior, record evidence that explains every consequential action, and enforce decisions with controls the model cannot argue with. Each article defines a layer's guarantee, its contract with its neighbours, what fails when it is skipped, and how to test that it holds, closing with conformance tests and a build order. Vendor-neutral and grounded in primary security and AI-security literature.
Every component of an AI agent can pass review while its behavior reaches an outcome nobody approved. Six layers, each with one guarantee, close that gap.
When each attempt is nearly free, a defence that usually works is a defence that eventually fails. Layer 1 turns that fact into rules every other layer must obey.
Every permission an agent holds can be justified on its own and still add up to reach nobody chose. Layer 2 makes that reach visible, bounded, and testable.
An agent cannot reliably tell instructions from data, so the design must. Layer 3 decides where every value came from before it may steer a privileged action.
Permissions are checked one call at a time, but harm is produced by sequences. Layer 4 approves and forbids combinations of actions as combinations.
A log says an action succeeded; an explanation says why it happened and who answers for it. Layer 5 makes every consequential action leave the second kind of record.
The model proposes; an independent mechanism decides, and returns the same verdict for the same facts however persuasively the request is worded.
An architecture is only as real as the evidence that it holds. This is how to adopt the six layers one outcome at a time, and how to prove each guarantee along the way.