A benchmark score is supposed to tell you what a model can do. This measurement series argues that for dangerous capabilities the number is only a lower bound, and a fragile one: the true capability can be hidden by a weak elicitation method, deliberately suppressed by a model that sandbags, or distorted by a model that notices it is being tested. Each article treats one threat to measurement validity — capability-versus-behavior, the elicitation gap, strategic under-performance, and evaluation awareness — with the data-scientist's tools of provenance, controls, and threats-to-validity, and closes with a defensible standard that pairs every capability claim with what it can and cannot prove. Product-agnostic and grounded in the primary dangerous-capability-evaluation and sandbagging literature.
2 series · 10 articles