Watch · narrated whiteboard episodesL3
A model that can tell it is being tested rather than deployed can, in principle, behave one way under evaluation and another in the wild — and the alignment literature now has proof-of-concept evidence that this test-versus-deploy divergence is real, persistent, and hard to train away. This paper series treats situational awareness and scheming as a measurement problem: when the subject of an evaluation can infer and act on whether it is observed, how do you build evaluations and safety cases that survive that divergence? The organizing contribution is a Test-Deploy Divergence protocol for probing behavior contingent on the model's belief about being watched. Product-agnostic, grounded in the primary situational-awareness, sleeper-agent, and in-context-scheming literature, and tied back to the AI-agent stack every time.