Skip to player
Research
Offensive AI Security
›
Red-Team Measurement
›
The Judge Problem: LLM-as-Judge Bias, Human-Label Reliability, and Calibrating a Grader
Contents
Host
Expert
Murali Chillakuru
Press play to begin the walkthrough.
0:00
16:42
1×
1.25×
1.5×
0.85×
CC
Red-Team Measurement · 2 / 5
The Judge Problem: LLM-as-Judge Bias, Human-Label Reliability, and Calibrating a Grader
Play
Back