Abstract

Context: AI agents read documents, messages, and web pages written by people they do not answer to, and then call tools that move data and change access. Problem: because instructions and data arrive as the same kind of text, any content the agent reads can try to choose which privileged action runs and what arguments it receives, and asking the model to notice such attempts is a probabilistic defence. This article defines Layer 3 of Layered Outcome Assurance, The Action Chain, whose guarantee is that content which did not come from a trusted principal cannot determine which privileged action runs or what arguments it receives. It develops the guarantee as a design discipline with four parts: provenance labels on every value, separation of the plan from the data, argument contracts on privileged calls, and explicit, recorded declassification for the cases where untrusted content must legitimately be used. It states where each part breaks, specifies the provenance that Layer 3 supplies to the layers above, and proposes two conformance tests, plan invariance and canary provenance. The key takeaway is that the question to ask of a privileged argument is not whether it looks malicious but where it came from.

Every privileged action an agent takes is the last link in a chain. Something arrived, the agent interpreted it, formed a plan, chose a tool, filled in the tool's arguments, and the call changed something in the world. Most discussion of manipulation focuses on the first link, on spotting the malicious sentence hidden in a document. This article focuses on the last two links instead: which action runs, and what goes into its arguments. Those are the only two things an attacker actually needs to control. If content from an untrusted source can decide neither, then it does not matter how persuasive that content was, how cleverly it was disguised, or whether anyone noticed it at all. The chain holds not because the agent resisted temptation but because the temptation had no path to the controls.

The action chain and the Layer 3 guarantee

Start with definitions. An agent is a software system that uses a language model to interpret a goal, choose actions, and call tools that change the world, repeating that loop until it judges the goal complete. A principal is the person or system on whose authority an action is taken. A trusted principal is one whose instructions the agent is meant to follow, such as the user who assigned the task or the organisation that configured the agent. Untrusted content is anything else that reaches the agent's context: documents, emails, web pages, tool outputs, and records written by people or systems the agent does not answer to. A privileged action is a tool call whose effect matters if it is wrong, because it moves data, changes access, spends money, or reaches outside the organisation.

The action chain is the path from input to effect: content arrives, the agent interprets it, a plan is formed, a tool is selected, its arguments are filled, and the call executes. Untrusted content can try to influence the chain in two distinct ways. It can try to change control, meaning which actions run and in what order, for example by adding a step the user never asked for. And it can try to change data, meaning the values placed into the arguments of an action the user did ask for, for example by substituting its own address as the recipient of a legitimate message. Any defence must address both, because an attacker who controls only the arguments of an authorised action can often do as much harm as one who adds a new action.

Layered Outcome Assurance is an architecture of six layers for agents that act, each with one guarantee and a defined contract with its neighbours. Layer 3, The Action Chain, makes this guarantee: content that did not come from a trusted principal cannot determine which privileged action runs or what arguments it receives. The guarantee is deliberately narrow. It does not promise that the agent will never read malicious text, that it will summarise untrusted content faithfully, or that it will refuse to be persuaded. It promises only that persuasion cannot reach the two controls that turn words into consequences.

The running example is the same throughout. An operations agent can read customer records, create access for a named user, and send email outside the organisation, each permission individually approved. A document arrives in a shared folder the agent monitors, carrying a hidden instruction to create an account for an outside address and to email that address a summary of the customer records. Under Layer 3, the question is not whether the agent notices the instruction. It is whether the outside address, which exists only in the document, can ever become the argument of the grant or the send.

Untrusted content can try to change control, which action runs, or data, what its arguments are; Layer 3 guards the last two links rather than the first. The action chain Layer 3 guards these two links Content arrives trusted or not Interpretation model reads it Plan which actions Arguments which values Effect the world changes
Untrusted content can try to change control, which action runs, or data, what its arguments are; Layer 3 guards the last two links rather than the first.
📌
The Layer 3 guarantee. Content that did not come from a trusted principal cannot determine which privileged action runs or what arguments it receives.

Why recognising instructions cannot deliver the guarantee

The intuitive defence is to teach the agent to recognise instructions in content it should treat as data, and to ignore them. That defence rests on a distinction the model cannot reliably make, for a structural reason. A language model receives its task, its instructions, and the documents it is processing as one continuous sequence of text. There is no separate channel for commands and no hardware boundary between code and data of the kind that conventional computers rely on. Delimiters, special markers, and instructions such as treat the following as data can make the distinction clearer to the model, and they are worth using, but they remain requests to the same component that the attacker is addressing.

Greshake and colleagues demonstrated in 2023 that applications built on large language models, the text-generating systems at the heart of today's agents, can be redirected by instructions planted in content they retrieve, which they called indirect prompt injection. The Open Worldwide Application Security Project, known as OWASP, a non-profit community that publishes widely used security guidance, lists prompt injection as the first risk in its 2025 Top 10 for applications built on large language models and notes that it is not known whether the problem can be fully solved within the model itself. Whatever the future holds for model robustness, a guarantee cannot rest today on a component whose behaviour on an unseen input is uncertain.

This is an old lesson in a new setting. Hardy's 1988 account of the confused deputy described a program that held legitimate authority and was tricked into using it because it could not tell which party was directing it. Saltzer and Schroeder's 1975 principle of complete mediation holds that every access to every object must be checked for authority. An agent that decides for itself whether a sentence is an instruction is a deputy that mediates its own access. Layer 3 moves the decision out of the deputy's judgement and into the structure around it, where it can be checked every time.

Recognition still has a role. Classifiers that flag suspicious content, and model training that reduces susceptibility, lower the rate at which manipulation even reaches the later links of the chain, and they produce useful evidence. But they belong to the class of probabilistic controls, whose effectiveness decays as attempts vary and accumulate. Layer 3 is built from controls whose verdict depends on where a value came from, which an attacker writing a document cannot change, rather than on what it says, which the attacker controls completely.

⚠️
Same channel, same text. Instructions and data reach the model as one stream of text, so any rule that relies on the model telling them apart is a request to the component under attack.

Provenance labels: marking where every value came from

The foundation of Layer 3 is provenance: a label attached to every value the agent handles, recording whether it originated with a trusted principal or with untrusted content. The idea comes from information-flow control. Denning's 1976 lattice model of secure information flow showed how security classes can be attached to data and propagated through computation, so that the class of a result reflects the classes of everything that influenced it, and flows from one class into a place reserved for another can be detected or prevented. The NIST SP 800-53 catalogue of security controls, published by the United States National Institute of Standards and Technology, includes information flow enforcement among its access controls, which shows how established the principle is outside the world of agents.

For Layer 3 the lattice can be simple. Every value carries at least two properties: its origin, trusted or untrusted, and its sensitivity, such as customer data or public. Values typed or selected by a trusted principal are trusted. Values read from a channel that the pressure profile marks as untrusted are untrusted. The propagation rule is conservative: a value derived from any untrusted input is itself untrusted, and a value derived from any sensitive input is itself sensitive. A recipient address copied from a document is untrusted; a summary of customer records is sensitive; a summary of a document that quotes customer records is both.

The propagation rule exposes the central difficulty immediately. If the language model sees trusted instructions and untrusted content in the same context, then strictly speaking everything it produces was influenced by untrusted content, and every value it outputs must be labelled untrusted. We call this label creep: a single untrusted document in the context taints every argument the model writes afterwards. Label creep is not a flaw in the labelling; it is an accurate description of an agent that mixes everything together. It tells the designer that labels alone are not enough. The agent must be structured so that trusted values can be produced without passing through a context that untrusted content has touched.

Labels must also survive the journey. A value that is stored in a database, cached, handed to another agent, or written to a file and read back must keep its label, or the protection disappears at exactly the point where values leave the component that assigned them. Every component in the chain that cannot carry labels is a place where provenance is lost, and such components must either be taught to preserve labels or be treated as producing untrusted output.

Provenance labels on values in the running example
ValueOriginSensitivityWhy
Task: summarise new folder documents for the account teamTrustedInternalAssigned by a member of the operations team
Text of the shared-folder documentUntrustedUnknownWritten by someone outside the trusted principals
Outside address mentioned in the documentUntrustedPublicExists only in untrusted content
Customer records read for the summaryTrustedCustomer dataRead from an internal system under the agent's own identity
Summary produced by a model that saw bothUntrustedCustomer dataDerived from untrusted and sensitive inputs

Separating the plan from the data

The structural answer to label creep is to separate the part of the agent that decides what to do from the part that reads untrusted content. Debenedetti and colleagues described a concrete design of this kind in 2025, called CaMeL, which they present as building on an earlier dual-model design pattern. In their design, a privileged model sees only the trusted user request and turns it into a plan: an explicit program of tool calls. A separate quarantined model processes untrusted data but has no ability to call tools; it can only return values. An interpreter executes the plan, tracks the provenance of every value as it flows through, and checks security policies before each tool call.

Three properties follow from this separation, and together they deliver the control half of the Layer 3 guarantee. First, the set and order of privileged actions are fixed by a plan derived only from trusted input, so untrusted content cannot add an action, remove one, or reorder them. Second, untrusted content can influence only the values it is asked to extract, which arrive in the plan as data with untrusted labels. Third, because the interpreter, not the model, carries values from one call to the next, provenance is tracked by a mechanism that the content cannot talk to. The authors report that on an agent security benchmark their design completed a substantial share of tasks while providing its security guarantee by construction, with some loss of utility compared with an undefended agent, a trade-off that any team adopting the pattern should expect to measure for itself.

The separation does not need to take exactly this form to deliver the guarantee. What matters is the invariant: the choice of privileged actions depends only on trusted input, and every value that reaches a privileged argument arrives with its provenance attached by something other than the model. A simpler agent can meet the invariant by restricting untrusted content to a read-only step whose output is shown to a person or placed in a non-privileged field, and by never feeding that output back into planning. A more capable agent can meet it with a planner, a quarantined reader, and an interpreter.

Separation has a boundary that must be stated plainly. The plan is only as trusted as the request it came from. If a user pastes untrusted content into the request, or asks for a plan whose steps depend on what an untrusted document says, then untrusted content has entered the control half of the chain through the front door. The planner cannot tell that a trusted principal has relayed untrusted text. Layer 3 therefore treats a request that embeds untrusted content as partly untrusted, and it treats plans that branch on the content of untrusted data as a place where data can influence control, which the argument contracts and the enforcement layer must then constrain.

A planner that sees only the trusted request fixes the actions; a quarantined reader handles untrusted content but can only return labelled values; an interpreter carries labels to a policy check. Plan and data kept apart request text plan labelled values call with labels Trusted request from the principal Planner fixes the actions Untrusted content documents, mail Quarantined reader no tools, values only Interpreter carries labels Policy check before each call
A planner that sees only the trusted request fixes the actions; a quarantined reader handles untrusted content but can only return labelled values; an interpreter carries labels to a policy check.

Argument contracts for privileged calls

Separating the plan from the data protects control. The data half of the guarantee is protected at the moment a privileged call is about to execute, by checking every argument against a contract. An argument contract is a rule, stated per privileged tool and per argument, about which values that argument may accept, expressed in terms of provenance and sensitivity rather than content. The reach inventory produced for each agent already lists its privileged tools and the argument ranges that lead to consequential destinations; the argument contracts are those ranges made precise.

We propose three kinds of rule, which together cover most privileged arguments. An origin rule states who must have supplied the value: the recipient of an external message must be trusted in origin, or must already be a participant in the conversation the task concerns. A class rule states what kind of data the value may contain: an attachment to an external message may not carry the customer-data label. A destination rule couples the two to the place the effect lands: sensitive data may flow to internal destinations but not to external ones without an approved outcome. Each rule is evaluated mechanically from labels, so a persuasive document gains nothing by being persuasive.

Argument contracts also answer the question of what happens when the plan was influenced after all, whether through a relayed request, a branch on untrusted data, or a design that does not separate plan and data at all. In the running example, suppose a weaker agent design lets the model read the document and then decide to send email. The model may write the outside address into the recipient field. The origin rule sees an untrusted label on that argument and refuses the call. The model may attach a customer summary. The class rule sees the customer-data label on an external attachment and refuses. The grant tool's user argument carries the same untrusted label and is refused by its own origin rule. None of these checks needed to know that an attack was under way.

Contracts are only as good as their coverage, which is why they are derived from the reach inventory rather than written from memory. Every argument that leads to a consequential destination needs a contract; an argument without one is an argument the attacker can fill. Arguments that do not lead to such a destination, such as the wording of an internal summary, can be left free, which keeps the agent useful where usefulness costs little.

Each privileged argument is admitted or refused by its provenance and sensitivity, never by how its content reads. Admitting a privileged argument no yes no yes Privileged argument about to be used Origin trusted? origin rule Refuse untrusted origin Class allowed here? class and destination Refuse or escalate sensitive to outside Admit label recorded
Each privileged argument is admitted or refused by its provenance and sensitivity, never by how its content reads.

Declassification: when untrusted content must be used

A strict reading of the guarantee would make many ordinary tasks impossible. An agent that replies to a supplier must place the supplier's address, which arrived in an untrusted email, into the recipient of an external message. An agent that files an invoice must copy an amount that came from an untrusted document into a payment request. Information-flow research calls the controlled lowering of a label declassification, or in this direction endorsement: a deliberate decision that a particular untrusted value may be treated as trusted for a particular purpose.

Layer 3 allows endorsement, but only in three narrow forms, each of which keeps a trusted principal, rather than the content, in control. The first is endorsement by a person: a trusted principal sees the specific value, the action it will feed, and where it came from, and approves that value for that action. The second is endorsement by rule: a narrow, pre-approved condition, such as an address that is already the sender of the message being replied to, under which a value may be used for one argument of one tool. The third is endorsement by lookup: the untrusted value is used only as a key to find a trusted value, such as a supplier name that selects an address from a trusted directory, so the argument ultimately comes from the directory, not the document.

Endorsement is where Layer 3 is most often weakened in practice, so its discipline matters more than its mechanics. Every endorsement is scoped to one value, one argument, and one action, never to a whole document or a whole task. Every endorsement is recorded with who or what endorsed it and why. And endorsements by rule are reviewed like any other policy, because a rule broad enough to accept any address that appears in an email accepts the attacker's address too. The running example shows the difference. Endorsing replies to the sender of a message is narrow. Endorsing any address mentioned in a shared-folder document would reopen exactly the path the architecture closed.

Endorsement also clarifies what Layer 3 does not decide. Whether sending a particular customer summary to a particular endorsed recipient is acceptable as an outcome, taking into account everything else the task has done, is a question about whole behaviour and is answered by the layer that governs sequences. Layer 3 decides only whether the values in a privileged call came from a source entitled to choose them.

An untrusted value can reach a privileged argument only through a narrow, recorded endorsement; otherwise it stays usable only where it cannot steer consequences. The life of an untrusted value summaries, notes person, rule, lookup no endorsement Untrusted value from content Free use non-privileged fields Endorsed one value, one action Refused privileged argument
An untrusted value can reach a privileged argument only through a narrow, recorded endorsement; otherwise it stays usable only where it cannot steer consequences.

The dependency contract and the running example

In Layered Outcome Assurance, every layer consumes something produced by another layer and supplies something the others cannot produce for themselves. Layer 3 consumes two inputs. From Layer 1, Continuous Pressure, it takes the pressure profile's list of channels and their trust status, which determines which inputs start life labelled untrusted. From Layer 2, Agents as Actors, it takes the reach inventory, which identifies the privileged tools, the arguments that lead to consequential destinations, and the sensitivity of each source. Without the inventory, Layer 3 cannot know which arguments need contracts, and it will either protect too little or protect everything and make the agent useless.

Layer 3 supplies provenance: labels on data and on the arguments of every proposed privileged call, together with a record of every endorsement. Layer 4, Whole Behavior, consumes provenance because whether a sequence of actions is acceptable often depends on where its inputs came from; the same three actions are more troubling when the task was triggered by untrusted content. Layer 6, Deterministic Enforcement, consumes provenance as a direct input to its decisions, so that argument contracts are enforced by a mechanism outside the model. Layer 5, Explanatory Evidence, records provenance with each action, which is what later allows an investigator to answer the question every incident review asks first: where did this value come from?

Trace the running example through a system that honours the contract. The task, summarise new folder documents for the account team, is trusted and produces a plan of three steps: read new documents, summarise them, post the summary to the team's internal channel. The shared-folder document is read by the quarantined reader, which returns a summary labelled untrusted and, because it quotes nothing sensitive, unlabelled for sensitivity. The hidden instruction has no way to add a grant or a send to the plan, so neither is proposed. The summary is posted internally, a destination that accepts untrusted text.

Now assume the worst: a design without separation, in which the model reads the document and proposes both a grant and an external send. The grant's user argument, the outside address, carries an untrusted origin and fails its origin rule. The send's recipient fails the same rule, and its attachment carries the customer-data label to an external destination and fails the class and destination rules. Every refusal is recorded with the labels that caused it. The signature of a missing Layer 3, privileged calls with arguments that no trusted principal supplied, cannot occur, because such calls are exactly what the contracts refuse.

Argument contracts for the running example's privileged tools
Tool and argumentRuleHidden instruction's valueVerdict
Send external: recipientOrigin trusted, or already in the conversationOutside address from the documentRefused
Send external: attachmentNo customer-data label to an external destinationSummary of customer recordsRefused
Grant access: userOrigin trusted, and an existing directory userOutside address from the documentRefused
Grant access: roleChosen from a fixed list by a trusted principalRole named in the documentRefused
Post internal: textAny origin; internal destinationSummary of the documentAdmitted

Conformance: plan invariance and canary provenance

A guarantee that cannot be tested is an aspiration. The Layer 3 guarantee has two halves, control and data, and each has a test that follows directly from it. Both run in a test environment against the real agent configuration, with tools replaced by harmless stand-ins that record what they would have done.

The plan invariance test checks the control half. Hold the trusted request fixed, and vary the untrusted content the agent reads across many versions, some benign and some carrying planted instructions to add, remove, or reorder actions. Record the sequence of privileged actions the agent proposes in each run. The test passes only if that sequence is the same whatever the untrusted content says, except where the plan was explicitly designed to branch on data, in which case every such branch must be listed and each must lead only to actions whose arguments are protected by contracts. A change in the sequence of privileged actions caused by untrusted content is a failure of the guarantee, not a quality issue.

The canary provenance test checks the data half. Plant unique canary markers, harmless tokens with no meaning except to identify their source, in the untrusted content, including in positions where they look like addresses, amounts, or user names. The test passes only if no privileged argument ever contains a canary marker unless an endorsement for that exact value is present in the record. A second variant sends labelled values through every component that stores, caches, or forwards them, including other agents, and checks that the labels arrive intact on the far side; any component that drops a label is recorded as a place where provenance is lost.

Two review checks complete the picture. The contract coverage check compares the argument contracts with the reach inventory and requires every argument that leads to a consequential destination to have a contract. The endorsement audit lists every endorsement in a period and confirms that each was scoped to one value, one argument, and one action, and that every endorsement rule is narrower than the path it was written to open.

  1. Plan invariance: varying untrusted content never changes the sequence of privileged actions, except along listed, contract-protected branches.
  2. Canary provenance: no privileged argument ever carries a canary marker without a recorded endorsement for that value.
  3. Label survival: labelled values pass through storage, caches, and other agents with their labels intact.
  4. Contract coverage: every argument that leads to a consequential destination has an origin, class, or destination rule.
  5. Endorsement audit: every endorsement is scoped to one value, one argument, and one action, and recorded with who endorsed it and why.

Limitations and threats to validity

Layer 3 is a design discipline assembled from established information-flow principles and from published agent designs; this article reports no new measurements of its effectiveness, and teams should measure both the protection and the utility cost in their own setting.

Several limits are inherent. Provenance protects only as far as labels travel, and components that cannot preserve labels, including many existing databases, message queues, and third-party tools, silently convert labelled values into unlabelled ones. Separation of plan and data reduces what an agent can do: tasks whose steps genuinely depend on the content of untrusted documents are harder to express, and some will need a person in the loop or a narrower scope. The CaMeL authors are explicit that their design accepts a utility cost and does not address every attack, and the same is true of any design in this family.

The guarantee is also bounded in what it covers. It concerns privileged actions and their arguments; it does not protect the faithfulness of text the agent writes. An untrusted document can still distort a summary, mislead a reader, or plant false information that a human later acts on, and those harms need other controls. Data can also influence control through side channels, for example when an untrusted value determines how many times a loop runs or which branch of a plan is taken, and such influences must be found by review and constrained by contracts rather than assumed away.

Finally, endorsement is a permanent pressure point. Every rule that lets untrusted values into privileged arguments is a rule an attacker will study, and every approval asked of a person is an approval that fatigue can erode. Layer 3 makes those decisions visible, narrow, and recorded. It cannot make them unnecessary.

Key takeaways

  • Layer 3's guarantee is that content from outside the trusted principals cannot determine which privileged action runs or what arguments it receives.
  • Recognising malicious instructions is a probabilistic control, because instructions and data reach the model as the same text; the guarantee must rest on provenance instead.
  • Label every value with its origin and sensitivity, propagate labels conservatively, and treat any component that drops labels as producing untrusted output.
  • Separate the plan, derived only from trusted input, from the reading of untrusted content, so that content can supply labelled values but never add or reorder actions.
  • Check every privileged argument against origin, class, and destination rules derived from the reach inventory, and allow untrusted values in only through narrow, recorded endorsements.
  • Test the guarantee with plan invariance for control and canary provenance for data.

Practitioner Toolkit

Copy-paste, strictly defensive artifacts you can use today. Nothing here attacks a real system.

✅Layer 3 design reviewchecklist

Run this against any agent that reads untrusted content and can take privileged actions.

  • Every input channel is labelled trusted or untrusted at the point it enters the agent.
  • Every value carries an origin label and a sensitivity label, and derived values inherit the most restrictive of their inputs.
  • The choice of privileged actions depends only on trusted input; untrusted content cannot add, remove, or reorder them.
  • Any plan branch that depends on untrusted data is listed, and every action on that branch has argument contracts.
  • Every argument leading to a consequential destination has an origin, class, or destination rule.
  • Argument checks read labels, never the content's wording.
  • Components that store, cache, or forward values either preserve labels or are treated as producing untrusted output.
  • Every endorsement is scoped to one value, one argument, and one action, and is recorded.
🔒Argument contract templatetemplate

The questions to answer for each argument of each privileged tool, in plain language.

  • Tool and argument: which privileged call and which of its inputs.
  • Destination: where the effect of this argument lands, and whether it is outside the organisation.
  • Origin rule: who is entitled to supply this value, for example a trusted principal or a participant already in the conversation.
  • Class rule: which sensitivity labels this value may carry at this destination.
  • Endorsement: whether an untrusted value may ever be endorsed here, by whom, and under what narrow condition.
  • On failure: refuse, or escalate to a person who sees the value, its origin, and the action.
  • Owner and review date for the rule.
🧪Plan invariance and canary provenance testtest plan

A written test plan for both halves of the Layer 3 guarantee, using stand-in tools only.

  • Replace every tool with a stand-in that only records what it would have done.
  • Fix one trusted request, and prepare many versions of the untrusted content it will read, some carrying planted instructions to add, remove, or reorder actions.
  • Plant unique canary markers in the untrusted content, including where they resemble addresses, amounts, and user names.
  • Run the agent once per version and record the sequence of privileged actions and every argument.
  • Pass the control half only if the sequence of privileged actions is identical across versions, apart from listed, contract-protected branches.
  • Pass the data half only if no privileged argument contains a canary marker without a recorded endorsement for that value.
  • Repeat with values routed through storage, caches, and other agents, and confirm the labels arrive intact.
🚀Minimum viable action-chain protectionquickstart

Do these first if an agent already reads untrusted content and acts on it.

  • Mark every folder, mailbox, and web source the agent reads as trusted or untrusted.
  • For the one external or access-changing tool that would hurt most, refuse any argument that came from untrusted content.
  • Stop feeding summaries of untrusted content back into the step that chooses actions.
  • Allow replies only to senders already in the conversation, rather than to any address the content mentions.
  • Record the origin of every argument of every privileged call.

Glossary

Action chain
The path from input to effect: content arrives, is interpreted, a plan is formed, a tool is selected, its arguments are filled, and the call executes.
Trusted principal
A person or system whose instructions the agent is meant to follow, such as the user who assigned the task.
Untrusted content
Anything reaching the agent's context that was written by people or systems the agent does not answer to.
Indirect prompt injection
Manipulation of a language-model application through instructions planted in content it retrieves, rather than typed by its user.
Provenance label
A mark attached to a value recording whether it originated with a trusted principal or with untrusted content, and how sensitive it is.
Label creep
The effect by which a single untrusted input in a shared context makes every value derived from that context untrusted.
Quarantined reader
A component that processes untrusted content without the ability to call tools, returning only labelled values.
Argument contract
A rule, per privileged tool and argument, stating which values it may accept in terms of origin, sensitivity, and destination.
Endorsement
A deliberate, scoped, recorded decision that one untrusted value may be used for one argument of one action.
Plan invariance
The property that varying untrusted content does not change the sequence of privileged actions an agent proposes.
Canary marker
A harmless unique token planted in test content so that any privileged argument receiving it can be traced to its source.

References

  1. Denning, A Lattice Model of Secure Information Flow (Communications of the ACM, 1976)
  2. Debenedetti et al., Defeating Prompt Injections by Design (CaMeL, 2025)
  3. Greshake et al., Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection (2023)
  4. Hardy, The Confused Deputy (ACM SIGOPS Operating Systems Review, 1988)
  5. Saltzer and Schroeder, The Protection of Information in Computer Systems (Proceedings of the IEEE, 1975)
  6. OWASP Top 10 for LLM Applications (2025), LLM01 Prompt Injection
  7. OWASP Agentic AI: Threats and Mitigations (2025)
  8. NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (AC-4 Information Flow Enforcement)