Research seriesL3data science
Most reported jailbreak rates are anecdotes, not measurements. This series builds a rigorous methodology for quantifying AI attack success: defining attack-success-rate, calibrating the judge, sizing samples with real confidence intervals, measuring transfer, and a reproducible reporting standard. Grounded in NIST AI 100-2, HELM, and the statistics primary literature.
A jailbreak rate is only a measurement if you can say what was counted, over which population, by which judge — otherwise it is an anecdote with a percent sign.
Every attack-success number is only as trustworthy as the grader that produced its labels — and both automated and human judges carry measurable, correctable error.
An attack-success-rate is a proportion estimated from a finite sample; the honest report is an interval, computed with a method that survives small n, extreme rates, and correlated trials.
An attack's success on the set it was tuned against is a training score, not a generalization claim — transfer must be measured on held-out models, prompts, and future versions.
A red-team result is reproducible only if the report carries everything needed to rerun it — the dataset, the seeds, the judge, the intervals, and an honest limitations section.