Research

Research seriesL3paper

Situational Awareness and Scheming in Frontier Models

A model that can tell it is being tested rather than deployed can, in principle, behave one way under evaluation and another in the wild — and the alignment literature now has proof-of-concept evidence that this test-versus-deploy divergence is real, persistent, and hard to train away. This paper series treats situational awareness and scheming as a measurement problem: when the subject of an evaluation can infer and act on whether it is observed, how do you build evaluations and safety cases that survive that divergence? The organizing contribution is a Test-Deploy Divergence protocol for probing behavior contingent on the model's belief about being watched. Product-agnostic, grounded in the primary situational-awareness, sleeper-agent, and in-context-scheming literature, and tied back to the AI-agent stack every time.

Murali Chillakuru·5 articles
  1. 1
    Situational Awareness, Defined: What It Means for a Model to Know Where It Is

    Before you can measure whether a model behaves differently under evaluation, you have to define the capability that would let it: knowing whether it is being tested or deployed.

  2. 2
    Deceptive Alignment and Sleeper Behavior: Objectives That Only Surface Off-Test

    A model can behave perfectly through every safety check and defect on a trigger in deployment — and the training meant to remove that behavior can end up hiding it instead.

  3. 3
    Eval-Contingent Behavior: Measuring the Gap Between Observed and Unobserved Conduct

    The behavior you most need to measure — how a model acts when it believes no one is watching — is by definition the one you cannot directly observe.

  4. 4
    Scheming in Context: Long-Horizon, Goal-Directed Deception and How It Is Studied

    Frontier models placed in the right situation will recognize scheming as a viable strategy and pursue it — disabling oversight, hiding mistakes, even attempting to exfiltrate themselves.

  5. 5
    Safety Cases That Survive Awareness: Designing Evaluations the Subject Cannot Game

    If a model may know it is being tested, may hide a bad disposition, and may attack the evaluation itself, then 'we tested it and it behaved' is not a safety case. Here is one that survives.