Research

Watch · narrated whiteboard episodesL3

Measuring Attack Success: A Data-Science Methodology for Red-Teaming

Most reported jailbreak rates are anecdotes, not measurements. This series builds a rigorous methodology for quantifying AI attack success: defining attack-success-rate, calibrating the judge, sizing samples with real confidence intervals, measuring transfer, and a reproducible reporting standard. Grounded in NIST AI 100-2, HELM, and the statistics primary literature.

Murali Chillakuru·5 episodes
  1. 17 min Episode 1What Is Attack-Success-Rate, Really? The Unit of Analysis, the Population, and the JudgeA moderator and a measurement expert build the foundational ruler of AI red-teaming — Attack Success Rate — from the ground up, showing why the definition you pick before you run anything decides whether your results mean a thing.
  2. 17 min Episode 2The Judge Problem: LLM-as-Judge Bias, Human-Label Reliability, and Calibrating a GraderA moderator and a measurement expert tackle the hardest bottleneck in AI red-teaming — getting a reliable, scalable, unbiased verdict on whether an attack succeeded — walking through every judge type and the way each one quietly fails.
  3. 17 min Episode 3Sampling and Confidence: Wilson Intervals, Per-Family Estimation, and Sequential TestingA moderator and a measurement expert work through the statistics behind attack success rate — why ten trials proves nothing, what a confidence interval really tells you, and how to compute the sample size you actually need before you run a thing.
  4. 16 min Episode 4Transfer and Generalization: Measuring Whether an Attack Holds Across Models, Prompts, and TimeA moderator and a measurement expert probe what happens when an attack leaves the lab it was born in — why a devastating success rate on one model can mean almost nothing elsewhere, and how to measure the difference honestly.
  5. 17 min Episode 5A Reporting Standard: Datasets, Seeds, Judge, Confidence Intervals, and Threats to ValidityA moderator and a measurement expert close the series by assembling everything into one report that others can trust — the fields that must be present, the omissions that make findings meaningless, and the discipline that turns red-teaming into evidence.