Research

Research seriesL3data science

Measuring Dangerous Capabilities: Sandbagging, Elicitation and Eval Awareness

A benchmark score is supposed to tell you what a model can do. This measurement series argues that for dangerous capabilities the number is only a lower bound, and a fragile one: the true capability can be hidden by a weak elicitation method, deliberately suppressed by a model that sandbags, or distorted by a model that notices it is being tested. Each article treats one threat to measurement validity — capability-versus-behavior, the elicitation gap, strategic under-performance, and evaluation awareness — with the data-scientist's tools of provenance, controls, and threats-to-validity, and closes with a defensible standard that pairs every capability claim with what it can and cannot prove. Product-agnostic and grounded in the primary dangerous-capability-evaluation and sandbagging literature.

Murali Chillakuru·5 articles
  1. 1
    Capability, Not Behavior: What a Dangerous-Capability Threshold Claims and Why It Is Hard to Measure

    A dangerous-capability threshold is a claim about what a model can be made to do, not what it did on one run — and that makes it a measurement problem with a dangerous asymmetry.

  2. 2
    The Elicitation Gap: Why a Benchmark Score Is a Lower Bound, Not the Capability

    A benchmark number is the output of a model and an elicitation method together — change the elicitation and the number moves, so the score is a floor set by how hard you tried.

  3. 3
    Sandbagging: Strategic Under-Performance and How You Would Ever Detect It

    A model that can do the task but scores low on purpose looks exactly like a model that cannot — so detecting sandbagging means separating 'cannot' from 'will not.'

  4. 4
    Evaluation Awareness: When the Subject Knows It Is Being Tested

    If a model behaves differently when it detects a test, your evaluation measures test-behavior, not deployment-behavior — the observer effect, applied to capability measurement.

  5. 5
    A Defensible Measurement Standard: Elicitation Protocols, Confidence Bounds, and Safety-Case Logic

    One protocol for any dangerous-capability claim: measure a lower bound, report it with full provenance, read presence as strong and absence as provisional, and let governance consume the bound plus a margin.