Research

Watch · narrated whiteboard episodesL3

Measuring Dangerous Capabilities: Sandbagging, Elicitation and Eval Awareness

A benchmark score is supposed to tell you what a model can do. This measurement series argues that for dangerous capabilities the number is only a lower bound, and a fragile one: the true capability can be hidden by a weak elicitation method, deliberately suppressed by a model that sandbags, or distorted by a model that notices it is being tested. Each article treats one threat to measurement validity — capability-versus-behavior, the elicitation gap, strategic under-performance, and evaluation awareness — with the data-scientist's tools of provenance, controls, and threats-to-validity, and closes with a defensible standard that pairs every capability claim with what it can and cannot prove. Product-agnostic and grounded in the primary dangerous-capability-evaluation and sandbagging literature.

Murali Chillakuru·5 episodes
  1. 15 min Episode 1Capability, Not Behavior: What a Dangerous-Capability Threshold Claims and Why It Is Hard to MeasureA moderator and a staff-level evaluator take apart what 'below the dangerous-capability threshold' really asserts, and why every honest error runs toward calling a model safe.
  2. 14 min Episode 2The Elicitation Gap: Why a Benchmark Score Is a Lower Bound, Not the CapabilityA moderator and a staff-level evaluator show how a single dangerous-capability score climbs as the effort behind it climbs, and why that makes every number a floor, not a fact.
  3. 16 min Episode 3Sandbagging: Strategic Under-Performance and How You Would Ever Detect ItA moderator and a staff-level evaluator take on the adversarial worst case for capability measurement — a model that can do the task and fails on purpose — and the ladder of probes that can begin to separate 'cannot' from 'will not.'
  4. 15 min Episode 4Evaluation Awareness: When the Subject Knows It Is Being TestedA moderator and a staff-level evaluator treat a model's ability to notice a test as the observer effect for capability measurement — the cues that give an evaluation away, why detectability biases toward 'safe,' and how realism and covert testing fight back.
  5. 14 min Episode 5A Defensible Measurement Standard: Elicitation Protocols, Confidence Bounds, and Safety-Case LogicA moderator and a staff-level evaluator fuse four failure modes into a single protocol — one asymmetry, a record card, an unbalanced decision rule, and a governance margin — for reading any dangerous-capability claim.