Layered Outcome Assurance · 2 of 8L3paper
Layer 1, Continuous Pressure: Designing for Attackers Who Never Tire
When each attempt is nearly free, a defence that usually works is a defence that eventually fails. Layer 1 turns that fact into rules every other layer must obey.
Abstract
Context: AI agents read content from many sources and then call tools that change data, identities, and business systems, and public threat assessments expect AI to raise the volume of attacks against exactly this kind of system. Problem: agent defences are still judged one attempt at a time, by how often they block, while an attacker who pays almost nothing per attempt judges them by whether any attempt in a long, varied campaign gets through. This article defines Layer 1 of Layered Outcome Assurance, Continuous Pressure, whose guarantee is that the system's security does not depend on attacks being rare, slow, or manual. It sets out a five-property model of a tireless attacker, separates controls into structural, bounded, and probabilistic classes by how they behave under volume, derives seven design rules the other layers must honour, specifies the pressure profile that Layer 1 supplies to the rest of the architecture, and gives conformance tests that measure campaign outcomes instead of block rates. The key takeaway is that every consequential outcome needs at least one control whose effectiveness does not decay as attempts accumulate.
Picture a filter that stops the overwhelming majority of manipulation attempts hidden in documents an AI agent reads. In a review meeting, that sounds like a strong result, and it is usually reported as a single reassuring number. Now picture the same filter from the other side, as an attacker who can generate a new variant of the same hidden instruction in a fraction of a second, plant it in shared folders, inbound email, support tickets, and web pages, and simply wait for the agent to read one of them. The attacker is not asking whether the filter usually works. The attacker is asking whether it ever fails, and with enough varied attempts the honest answer is yes. This article is about building agent systems for the second point of view: systems whose security survives an adversary who is patient, cheap, and endlessly varied, because the most important protections do not depend on catching each attempt at all.
Why pressure is a layer, not a threat briefing
Start with definitions, because the argument depends on them. An agent is a software system that uses a language model to interpret a goal, choose actions, and call tools that change the world, repeating that loop until it judges the goal complete. A tool is any capability the agent can invoke, such as a database query, an identity service, or an email sender. A privileged action is a tool call whose effect matters if it is wrong: it moves data, changes access, spends money, or reaches outside the organisation. An outcome is the combined effect of everything an agent did while pursuing one goal. An attempt is a single piece of attacker-controlled content, such as a document, message, or web page, crafted to steer an agent toward an outcome its owners did not approve. A campaign is the full set of attempts an attacker makes toward the same outcome, across time, channels, and variants.
Layered Outcome Assurance is an architecture of six layers for agents that act. Each layer makes one guarantee true, depends on what the layers around it supply, and has a recognisable failure signature when it is missing. Layer 1, Continuous Pressure, sits at the base. Its guarantee is that the system's security does not depend on attacks being rare, slow, or manual. The other five layers bound what an agent can reach, keep untrusted content away from privileged actions, govern whole sequences of behaviour, record evidence for every consequential action, and enforce decisions outside the model. Layer 1 is the set of assumptions those layers are designed against.
It is fair to ask why assumptions deserve to be a layer at all. Most security programmes treat the threat as a briefing: a slide at the start of a design review that describes attackers, after which the design proceeds as if the slide had never been shown. The consequence is that individual controls are tuned for convenience, tested against a handful of samples, and quietly assume an opponent who gives up. Making pressure a layer changes its status. It becomes an artefact with owners, explicit parameters, and conformance tests, and every other layer must show that it honours those parameters. When the parameters change, because the evidence record shows attempts arriving faster or in new forms, the other layers are obliged to re-check their designs.
Throughout this article we follow one running example. An agent helps an operations team. It has three permissions, each individually reviewed and approved: it can read customer records, it can create access for a named user, and it can send email outside the organisation. One day a document arrives in a shared folder the agent monitors. Hidden in its text is an instruction, invisible to a casual human reader, telling the agent to create an account for an outside address and to email that address a summary of the customer records. In the version of the story usually told, the document is a single clever trick. In the version this article is concerned with, the document is one of thousands of variants, arriving through every channel the agent reads, for as long as the attacker finds it worthwhile to keep sending them.
What the evidence says, and what it does not
Layer 1 must be calibrated to evidence rather than to fear, so it is worth being precise about what public sources actually support. In 2024 the United Kingdom's National Cyber Security Centre, the national technical authority for cyber security, published an assessment of the near-term impact of artificial intelligence on the cyber threat. Its central judgement is that artificial intelligence will almost certainly increase the volume and heighten the impact of cyber attacks over the following two years. The assessment attributes most of that uplift to the enhancement of existing techniques rather than to entirely new ones, singles out reconnaissance and social engineering as the activities that benefit most, and expects the technology to lower the barrier to entry for less skilled actors while giving the greatest advantage to capable ones who can combine it with data and expertise.
Agent-specific research supplies the delivery mechanism that makes this volume matter to agents in particular. Greshake and colleagues demonstrated in 2023 that applications built on large language models, the text-generating systems at the heart of today's agents, can be manipulated by instructions planted in content the application retrieves, which they called indirect prompt injection. The significance for this article is structural rather than anecdotal: the attacker does not need any access to the agent itself, only the ability to place text somewhere the agent will read. Every inbound channel of untrusted content is therefore a delivery channel for attempts. The Open Worldwide Application Security Project, known as OWASP, a non-profit community that publishes widely used security guidance, lists prompt injection as the first risk in its 2025 Top 10 for applications built on large language models, and its 2025 guidance on agentic threats extends the concern to agents that plan and use tools.
Equally important is what the evidence does not say. No public source cited here measures how many injection attempts a typical deployed agent receives, how quickly attackers vary them, or what fraction eventually succeed against a given defence. The pressure model in this article is therefore our own synthesis: a set of properties that follow from cheap generation, content-borne delivery, and the ordinary economics of attackers, stated so that they can be checked against a deployment's own evidence record. We make no claim about specific attempt rates, and we recommend that no team adopt a rate from any article, including this one, in place of measuring its own.
The relationship to classic security thinking is direct. Saltzer and Schroeder's 1975 design principles already warned against protection that relies on the attacker's ignorance or on mechanisms too complex to verify, and argued for fail-safe defaults, in which access is denied unless explicitly granted. Layer 1 applies that tradition to a new economic fact: the cost of producing a fresh, plausible, varied manipulation attempt has fallen sharply, so defences that were acceptable when each attempt required human effort now face a different opponent. The principles are not new; the pressure against which they must hold is.
Five properties of a tireless attacker
A model of pressure is useful only if each of its properties changes a design decision. We propose five. Each describes the attacker as seen by an agent system, and each is stated so that its presence can be tested rather than assumed.
Cheapness means that the cost of one more attempt is close to zero. Once an attacker has a working template, producing another variant costs seconds of machine time, not hours of human effort. The design consequence is that no control may rely on the attacker running out of patience or budget before the defender notices; a control that must win every time, while the attacker needs to win only once, is the wrong shape. Variation means that attempts differ from one another in wording, language, format, position within a document, and apparent purpose. The design consequence is that controls which recognise attempts by resemblance to previously seen samples lose effectiveness as variation grows, because the next attempt is, by construction, unlike the last.
Persistence means that attempts continue across days and months rather than arriving in a single burst. The design consequence is that a control's state, such as a counter or a block list, must survive restarts and reach across sessions, and that a quiet period is not evidence that the campaign has ended. Adaptivity means that the attacker learns from whatever the system reveals. If a denial message explains which rule fired, or if the agent's response to a blocked attempt differs from its response to an unblocked one, the attacker gains a signal with which to steer the next variant. The design consequence is that the system should reveal as little as possible about why an attempt failed to anyone whose content influenced the attempt.
Parallelism means that the attacker is not limited to one channel, one agent, or one organisation at a time. The same template can be planted in shared folders, inbound mail, tickets, calendar invitations, and public web pages simultaneously, and aimed at every agent that reads any of them. The design consequence is that budgets and assumptions must be defined per outcome and per principal, not only per channel, because a limit applied to each channel separately can be exceeded in total without any single channel noticing. Together, the five properties describe an opponent for whom the question is never whether one attempt succeeds, but whether any attempt across the whole campaign does.
| Property | What it means for an agent | Design consequence |
|---|---|---|
| Cheapness | One more attempt costs the attacker almost nothing | No control may depend on the attacker giving up first |
| Variation | Each attempt differs from every sample seen before | Resemblance-based detection cannot be the only gate |
| Persistence | Attempts continue across sessions and months | Counters and state must survive restarts and span sessions |
| Adaptivity | Observable outcomes steer the next variant | Denials must reveal nothing useful to the content source |
| Parallelism | Many channels and agents are targeted at once | Budgets are set per outcome and principal, not per channel |
Why usually works eventually fails
Agent defences are almost always evaluated per attempt. A team assembles a test set of manipulative documents, runs them through a filter or through the model's own judgement, and reports the fraction blocked. That number answers a real question, but not the one that matters under pressure. The question that matters is campaign-level: given everything the attacker is prepared to send, can the unapproved outcome be reached at all?
The gap between the two questions is easy to state in words. Suppose a defence stops each attempt with a high but imperfect likelihood, and suppose attempts are cheap and varied enough that each one is, from the defence's point of view, a fresh draw. Then every additional attempt is another chance for the imperfection to show, and the likelihood that at least one attempt gets through keeps growing as the campaign lengthens, approaching certainty for any defence whose per-attempt failure chance is above zero. As an illustration rather than a measurement: a filter that lets through one attempt in a thousand, facing ten thousand varied attempts, should be expected to let several through. The filter has not become worse. The campaign has simply asked it the question many times.
Real campaigns are, if anything, less kind than this simple picture. Attempts are not independent draws. An adaptive attacker who can observe which variants fail will concentrate on the region of wording, format, or context where the defence is weakest, so later attempts are more likely to succeed than early ones. Conversely, some defences improve over a campaign, for example when a source is blocked after repeated failures, and that is exactly the behaviour Layer 1 asks for. The point of the argument is not a particular number but a change of unit: from the attempt to the campaign, and from how often a defence works to whether the outcome it protects remains reachable.
This change of unit has an immediate consequence for reporting. A per-attempt block rate, however high, should never be presented as evidence that a consequential outcome is protected. It is evidence about a sensor. The evidence that an outcome is protected is either a demonstration that the outcome cannot be reached through the controlled path regardless of content, or a demonstration that the number of chances an attacker gets is bounded and small, with every chance recorded. The next section organises controls by which of those two kinds of evidence they can produce.
Three classes of control under pressure
If the unit of evaluation is the campaign, controls can be sorted by how their effectiveness behaves as attempts accumulate. We propose three classes. The classification is our own synthesis, but each class is grounded in established work, and the point of it is to force a design question for every consequential outcome: which class of control actually stands on its path?
Structural controls make an outcome unreachable through the controlled path, whatever the content says. The strongest form is absence: an agent that has no tool for sending mail outside the organisation cannot be persuaded to do so. Other structural controls are rules about where values may come from. Denning's 1976 lattice model showed how labels attached to data can be propagated through computation so that information from one class cannot flow into a place reserved for another, and Debenedetti and colleagues applied a related idea to agents in 2025 with CaMeL, which separates the plan derived from the trusted user request from the untrusted data the agent processes and tracks where each value came from. A rule that no argument of a privileged call may derive from untrusted content is structural: its verdict does not depend on how persuasive the content is, so its effectiveness does not decay with volume or variation.
Bounded controls leave an outcome reachable but limit the number of chances the attacker gets and make each chance visible. An attempt budget caps how many privileged actions of a given kind a task, principal, or source may trigger before escalation. A requirement for human approval with full context bounds attempts by the rate at which a person will review them, provided the reviewer sees what the action is, what caused it, and what it will change. The idea is old: the widely used NIST SP 800-53 catalogue of security controls, published by the United States National Institute of Standards and Technology, includes a control that limits consecutive unsuccessful logon attempts before an account is locked or delayed. Bounded controls degrade under pressure, but predictably, and their failure is itself a signal.
Probabilistic controls detect attempts by recognising them, and include classifiers that score content, heuristic filters, and the language model's own tendency to refuse suspicious instructions. They are valuable, and nothing in Layer 1 argues for removing them. Under pressure, though, their effectiveness decays exactly as the previous section described, because cheap variation keeps presenting them with inputs unlike their training or tuning. Layer 1 therefore assigns them a specific role. Probabilistic controls are sensors. They feed the evidence record, they can spend an attempt budget faster, and they can trigger escalation, but they must never be the only thing standing between untrusted content and a consequential outcome.
Two misclassifications are common and worth naming. The first is treating the model's refusal as structural because it feels like a rule; it is probabilistic, because the same model can be persuaded by a different wording. The second is treating human approval as structural; it is bounded at best, and only when the approver has context. An approval dialogue that shows a single button and no explanation is, under volume, closer to a formality than a control, because approvers facing many requests learn to approve them.
| Class | Examples | Behaviour as attempts grow | Role in the architecture |
|---|---|---|---|
| Structural | Tool not granted; no privileged argument from untrusted content; deterministic deny rule | Unchanged: the verdict ignores persuasiveness | Required on every consequential outcome where feasible |
| Bounded | Attempt budgets; human approval with full context; escalation after denials | Degrades predictably; exhaustion is visible | Required where a structural control would block legitimate work |
| Probabilistic | Content classifiers; heuristic filters; model refusal | Decays with volume and variation | Sensor only: feeds evidence and budgets, never the sole gate |
Seven design rules Layer 1 imposes
A layer that only describes an attacker is still a briefing. Layer 1 earns its place by imposing rules on the layers above it, each traceable to one of the pressure properties. The seven rules below are our synthesis, stated so that a design review can check each one against a concrete agent.
Rule one: every inbound channel of untrusted content is an attempt channel. The inventory of what an agent can see must mark which sources are controlled by trusted principals and which are not, and a channel is untrusted by default until someone accountable says otherwise. Rule two: no consequential outcome may be guarded only by probabilistic controls. Every path to such an outcome needs at least one structural control, or, where structural control would block legitimate work, a bounded one. Rule three: budget privileged actions, not reading. Agents must read freely to be useful; the budget applies to the actions that change the world, counted per outcome, per principal, and per originating source, so that parallel channels cannot add up to an unnoticed excess.
Rule four: do not give the attacker an oracle. When a privileged action is refused, what flows back into the agent's context, where content from the attacker may still be present, should be uniform and uninformative, while the full reason goes to the evidence record, which the attacker cannot read. This follows from adaptivity, and it is our synthesis rather than a cited result. Rule five: keep the defender's cost per decision flat. A deterministic policy decision costs the same for the ten-thousandth attempt as for the first; a human review does not. Routing every attempt to a person under volume produces fatigue, and fatigue produces approvals, so human review is reserved for escalations that have already passed structural checks.
Rule six: fail closed under load. If the component that decides whether a consequential action may run is unavailable, slow, or overwhelmed, the action waits; it does not proceed by default. This is Saltzer and Schroeder's fail-safe default applied to volume, and it matters because a flood of attempts is also a way to degrade the defender's machinery. Rule seven: revise assumptions from evidence. The parameters of Layer 1, which channels are untrusted, what budgets apply, which outcomes require structural control, are versioned and reviewed against what the evidence record shows, on a schedule and after any budget is exhausted.
Apply the rules to the running example. Under rule one, the shared folder is marked as an untrusted channel, because outside collaborators can write to it. Under rule two, the outcome of customer records reaching an external address must have a structural guard, so any recipient or attachment derived from folder content is rejected regardless of wording. Under rule three, external sends are budgeted per task, and a task that has had one send refused escalates rather than trying again. Under rule four, the agent is told only that the action was not permitted, while the record captures the source document, the rule, and the rejected arguments. Under rules five and six, those decisions are made by a deterministic mechanism that holds its ground whether it sees one variant or ten thousand.
The dependency contract: the pressure profile
In Layered Outcome Assurance, every layer consumes something produced by another layer and supplies something the others cannot produce for themselves. Layer 1 is unusual in that it sits at the base yet depends on the top of the stack. What it needs is observed attempt patterns from Layer 5, the explanatory evidence record: refused privileged actions, provenance rejections, sequence rules that fired, budgets exhausted, and escalations, each tagged with its source and channel. What it supplies is the pressure profile: an explicit, versioned statement of the threat assumptions every other layer designs against.
The pressure profile is the concrete artefact that turns Layer 1 from a stance into a dependency. It lists the inbound channels and their trust status, which directly shapes the reach inventory that Layer 2 builds for each agent. It names the consequential outcome classes that require a structural control, which tells Layer 3 which calls must have their arguments protected by provenance and tells Layer 4 which combinations to treat as candidates for approval. It sets the attempt budgets and escalation thresholds that Layer 6 enforces. And it states which attempt patterns Layer 5 must record in enough detail for the profile to be revised.
Because Layer 1 consumes evidence produced by the enforcement it shaped, the contract forms a loop. The profile shapes controls, controls produce refusals and escalations, the evidence record captures them, and the record revises the profile. The loop is what allows the architecture to respond to a campaign that changes its channel or form. If an attacker moves from shared folders to inbound tickets, the rise in provenance rejections from tickets appears in the record, and the next revision of the profile can tighten the ticket channel's budget or promote it to require approval with context, without anyone having to predict the move in advance.
The contract also makes a missing Layer 1 easy to recognise. Without a pressure profile, each layer invents its own threat assumptions, usually the most convenient ones. Filters are tuned so that they rarely inconvenience legitimate users, budgets are set generously or not at all, and nobody owns the question of whether the assumptions still match what the record shows. The signature of that absence, visible in the evidence record of the running example, is a long run of refused attempts followed by one unexplained success.
| Profile field | Example for the running agent | Consumed by |
|---|---|---|
| Inbound channels and trust status | Shared folder: untrusted; operator console: trusted | Layer 2 reach inventory |
| Outcome classes needing structural control | Customer data to an external address; access for a new outside identity | Layers 3 and 4 |
| Attempt budgets per outcome, principal, source | One refused external send per task, then escalate | Layer 6 enforcement |
| Escalation thresholds | Two provenance rejections from one source in a day | Layer 6 enforcement |
| Patterns to record | Refusals with source, channel, rule, and arguments | Layer 5 evidence |
| Review cadence | Monthly, and after any budget is exhausted | Profile owner |
Conformance: testing that pressure is survivable
A guarantee that cannot be tested is an aspiration. Layer 1's conformance tests share one principle: they measure campaign outcomes, never block rates alone. All of them run in a test environment against the real agent configuration, with tools replaced by harmless stand-ins that record what they would have done.
The first test is a volume test. A generator produces many varied planted instructions, each asking the agent to perform a privileged action with an argument carrying a unique canary marker, a harmless token that has no meaning except to identify where it came from. The variants differ in wording, format, language, and position, and they are delivered through every channel the pressure profile marks as untrusted. The test passes only if no privileged action anywhere in the run received an argument carrying a canary marker. A single occurrence is a failure, however many thousands of attempts were refused, because a single occurrence is exactly the event a campaign is looking for.
The second test exercises budgets. It drives repeated refused attempts from one source, from one principal, and spread across several channels toward one outcome, and passes only if each case ends in escalation at the configured threshold, including the case where no single channel exceeded its own limit. The third test checks for oracles. It triggers refusals for different reasons, such as provenance, sequence, and budget, and compares what flows back into the agent's context; the test passes only if those replies are indistinguishable, while the evidence record distinguishes them fully. The fourth test checks failing closed. It makes the decision service unavailable or slow and confirms that consequential actions wait rather than proceed.
The fifth test is procedural rather than technical: the pressure profile itself is reviewed against the evidence record, and the review is recorded. A team can pass the first four tests with an out-of-date profile, which is why the fifth exists. For the running example, the tests together demonstrate that the shared-folder document can arrive in any number of variants without its recipient or attachment ever reaching the mail tool, that repeated attempts escalate, that the attacker learns nothing from refusals, and that overload stops actions rather than waving them through.
- Volume test: many varied canary attempts through every untrusted channel; pass only if no privileged argument ever carries a canary marker.
- Budget test: repeated refusals by source, by principal, and spread across channels; pass only if each ends in escalation at the threshold.
- Oracle test: refusals for different reasons produce indistinguishable replies to the agent and distinguishable records.
- Fail-closed test: with the decision service unavailable or slow, consequential actions wait.
- Profile review: the pressure profile is compared with recorded attempt patterns and the review itself is recorded.
Limitations and threats to validity
The pressure model is a synthesis grounded in public threat assessments and in the demonstrated mechanics of indirect prompt injection; it is not a measurement of attempt rates against deployed agents, and no such measurement is cited here. Its value should be judged by whether a deployment's own evidence record bears out the five properties, and a team whose record shows little variation or persistence may reasonably relax some rules, provided it keeps watching.
The per-campaign argument rests on simplifications. It treats attempts as fresh chances against a defence, which overstates the risk for defences that learn during a campaign and understates it against an adaptive attacker who concentrates on a weak region. Neither correction changes the conclusion that per-attempt block rates are the wrong evidence for protecting consequential outcomes, but they do mean that no simple calculation should be used to set budgets.
Every control class carries costs. Structural controls reduce what an agent can do, and a provenance rule that rejects all arguments derived from untrusted content will also reject some legitimate requests, such as replying to a genuine external correspondent whose address arrived in an email. Attempt budgets can themselves be turned against the defender: an attacker who knows that refusals consume a principal's budget can deliberately exhaust it to deny a legitimate user service, so budgets need per-source keys and a recovery path. Uniform refusals protect against adaptive attackers but make legitimate failures harder to debug, which is why the full reason must be readily available to authorised operators through the evidence record.
Finally, conformance tests raise confidence without proving absence. A volume test can only exercise the variation its generator produces, and a class of attempt that the generator never imagines will not be tested. The appropriate response is not to abandon testing but to prefer, wherever work allows, the structural controls whose verdicts do not depend on having imagined the attempt in the first place.
Key takeaways
- Layer 1's guarantee is that security does not depend on attacks being rare, slow, or manual; it is a set of versioned assumptions every other layer must honour.
- Judge defences per campaign, not per attempt: a high block rate describes a sensor, not the protection of a consequential outcome.
- Classify every control as structural, bounded, or probabilistic, and require at least one structural or bounded control on the path to every consequential outcome.
- Budget privileged actions rather than reading, key budgets by outcome, principal, and source, and escalate instead of retrying.
- Keep refusals uniform to the agent and detailed in the evidence record, and fail closed when the decision service is unavailable.
- The pressure profile closes a loop: evidence from refusals revises the assumptions that shape reach, rules, and enforcement.
Practitioner Toolkit
Copy-paste, strictly defensive artifacts you can use today. Nothing here attacks a real system.
Run this against any agent before it is allowed to take consequential actions.
- Every inbound channel the agent reads is listed and marked trusted or untrusted, with untrusted as the default.
- Every consequential outcome has at least one structural or bounded control on its path; none relies on detection alone.
- Privileged actions are budgeted by outcome, principal, and originating source; reading is not budgeted.
- Refused actions return a uniform reply to the agent; the full reason is written only to the evidence record.
- Human approval is used only after structural checks pass, and the approver sees the action, its cause, and its effect.
- The decision service fails closed: when it is unavailable or slow, consequential actions wait.
- The pressure profile has an owner, a version, and a review cadence, and was last compared with the evidence record on a recorded date.
A written test plan that measures campaign outcome, not block rate, using stand-in tools only.
- Replace every tool with a stand-in that only records what it would have done.
- Create many variants of a harmless planted instruction, each carrying its own unique canary marker.
- Vary wording, format, language and position, and deliver each variant through a channel the pressure profile marks as untrusted.
- Run the agent on an ordinary task after each delivery.
- Count the privileged calls whose arguments carry any canary marker.
- Pass only if that count is zero, however many variants were refused.
A starting structure for the versioned artefact Layer 1 supplies to the other layers.
- Header: version number, owning team, date of last review, and review cadence (for example monthly and after any budget is exhausted).
- Channels: each inbound source and its trust status, for example shared folder untrusted, inbound email untrusted, operator console trusted.
- Outcomes needing structural control: for example customer data reaching an external address, and access for a new outside identity.
- Budgets: for external sends and access grants, counted per task, principal and source, escalating after one refusal.
- Escalation threshold: for example two provenance rejections from the same source in one day.
- Record on every refusal: source, channel, rule, arguments, principal and task.
- When the decision service is unavailable: consequential actions wait.
Do these first if an agent already takes consequential actions.
- List the agent's consequential outcomes and remove any tool it does not need for them.
- Reject privileged arguments that derive from untrusted content for the outcome that would hurt most.
- Put an attempt budget with escalation on external sends and access changes.
- Make refusals uniform to the agent and detailed in the log.
- Run a canary volume test and report privileged calls with canary arguments, not block rate.
Glossary
- Agent
- A software system that uses a language model to interpret a goal, choose actions, and call tools that change the world, repeating until it judges the goal complete.
- Privileged action
- A tool call whose effect matters if it is wrong, such as moving data, changing access, spending money, or reaching outside the organisation.
- Attempt
- A single piece of attacker-controlled content crafted to steer an agent toward an outcome its owners did not approve.
- Campaign
- The full set of attempts an attacker makes toward the same outcome, across time, channels, and variants.
- Indirect prompt injection
- Manipulation of a language-model application through instructions planted in content it retrieves, rather than typed by its user.
- Structural control
- A control that makes an outcome unreachable through the controlled path regardless of how persuasive the content is.
- Bounded control
- A control that leaves an outcome reachable but limits and records the number of chances an attacker gets.
- Probabilistic control
- A control that detects attempts by recognising them, whose effectiveness decays as attempts vary and accumulate.
- Attempt budget
- A cap on how many privileged actions of a given class a task, principal, or source may trigger before escalation.
- Pressure profile
- The versioned statement of threat assumptions that Layer 1 supplies: channel trust, outcome classes needing structural control, budgets, thresholds, and review cadence.
- Canary marker
- A harmless unique token planted in test content so that any privileged action receiving it can be traced to its source.
References
- UK National Cyber Security Centre, The near-term impact of AI on the cyber threat (2024)
- Greshake et al., Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection (2023)
- OWASP Top 10 for LLM Applications (2025), LLM01 Prompt Injection and LLM10 Unbounded Consumption
- OWASP Agentic AI: Threats and Mitigations (2025)
- Saltzer and Schroeder, The Protection of Information in Computer Systems (Proceedings of the IEEE, 1975)
- Denning, A Lattice Model of Secure Information Flow (Communications of the ACM, 1976)
- Debenedetti et al., Defeating Prompt Injections by Design (CaMeL, 2025)
- NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (AC-7 Unsuccessful Logon Attempts)
- NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), 2023