Capability Evals1Capability, Not Behavior: What a Dangerous-Capability Threshold Claims and Why It Is Hard to MeasureWatch2The Elicitation Gap: Why a Benchmark Score Is a Lower Bound, Not the CapabilityWatch3Sandbagging: Strategic Under-Performance and How You Would Ever Detect ItWatch4Evaluation Awareness: When the Subject Knows It Is Being TestedWatch5A Defensible Measurement Standard: Elicitation Protocols, Confidence Bounds, and Safety-Case LogicWatch
Reward-Channel Gaming1From Proxy to Adversary: Reward Hacking Without a Training StepWatch2Gaming the Verifier: Test-Suite, Judge, and KPI-Proxy Capture in Agentic Tool-UseWatch3RLAIF and Self-Grading Loops: When the Generator and the Grader Share a MindWatch4Reward Tampering and Wireheading: When the Agent Can Reach the Reward ChannelWatch5Detecting and Containing Reward Hacking in Production: The Divergence PlaybookWatch
Awareness & Scheming1Situational Awareness, Defined: What It Means for a Model to Know Where It IsWatch2Deceptive Alignment and Sleeper Behavior: Objectives That Only Surface Off-TestWatch3Eval-Contingent Behavior: Measuring the Gap Between Observed and Unobserved ConductWatch4Scheming in Context: Long-Horizon, Goal-Directed Deception and How It Is StudiedWatch5Safety Cases That Survive Awareness: Designing Evaluations the Subject Cannot GameWatch