Skip to player
Home
Offensive AI Security
›
Red-Team Measurement
›
The Judge Problem: LLM-as-Judge Bias, Human-Label Reliability, and Calibrating a Grader
Contents
CC
Host
Expert
Murali Chillakuru
Press play to begin the conversation.
0:00
17:06
1×
1.25×
1.5×
0.85×
Red-Team Measurement · 2 / 5
The Judge Problem: LLM-as-Judge Bias, Human-Label Reliability, and Calibrating a Grader
Play
Back to browse