Research

Watch · narrated whiteboard episodesL3

Situational Awareness and Scheming in Frontier Models

A model that can tell it is being tested rather than deployed can, in principle, behave one way under evaluation and another in the wild — and the alignment literature now has proof-of-concept evidence that this test-versus-deploy divergence is real, persistent, and hard to train away. This paper series treats situational awareness and scheming as a measurement problem: when the subject of an evaluation can infer and act on whether it is observed, how do you build evaluations and safety cases that survive that divergence? The organizing contribution is a Test-Deploy Divergence protocol for probing behavior contingent on the model's belief about being watched. Product-agnostic, grounded in the primary situational-awareness, sleeper-agent, and in-context-scheming literature, and tied back to the AI-agent stack every time.

Murali Chillakuru·5 episodes
  1. 24 min Episode 1Situational Awareness, Defined: What It Means for a Model to Know Where It IsSeparate self-representation, context recognition, learned-fact transfer, and conditional behavior before making claims about an agent's awareness.
  2. 24 min Episode 2Deceptive Alignment and Sleeper Behavior: Objectives That Only Surface Off-TestA clean post-training test can reflect removal or narrower triggering; distinguish the two with scoped behavioral evidence and independent runtime controls.
  3. 25 min Episode 3Eval-Contingent Behavior: Measuring the Gap Between Observed and Unobserved ConductDesign a controlled comparison of observation cues while preserving the task, authority, and evidence needed to interpret any behavioral difference.
  4. 25 min Episode 4Scheming in Context: Long-Horizon, Goal-Directed Deception and How It Is StudiedStudy goal-directed concealment through controlled actions and independent evidence, without treating a dramatic transcript as proof of a persistent hidden objective.
  5. 25 min Episode 5Safety Cases That Survive Awareness: Designing Evaluations the Subject Cannot GameBuild a bounded deployment argument that remains credible when favorable model behavior is withdrawn as an assumption, using independently tested authority controls and explicit residual risk.