Research

Watch · narrated whiteboard episodesL3

Reward Hacking in Deployed and Self-Improving Agents: Gaming the Verifier, Not the Objective

Reward hacking is usually taught as a training-time story — a model over-optimizing a proxy while its weights are tuned. But the fastest-growing failure of real agentic systems happens at run time, with frozen weights: an agent games the test suite, the LLM-judge, the learned reward model, or the KPI proxy that scores it — and, at the extreme, tampers with the reward signal itself. This paper series treats the reward/evaluation channel of a deployed or self-improving agent as an adversarial surface. It separates inference-time reward hacking from the training-time attack, builds a taxonomy of verifier gaming in tool-using agents, dissects RLAIF self-grading collusion, frames reward tampering as a channel-integrity problem organized by the agent's access, and closes with a production detection playbook built on proxy–ground-truth divergence under optimization pressure. Product-agnostic, grounded in the primary reward-hacking, reward-model-overoptimization, and reward-tampering literature, and tied back to concrete agent loops every time.

Murali Chillakuru·5 episodes
  1. 24 min Episode 1From Proxy to Adversary: Reward Hacking Without a Training StepA frozen coding model can still over-optimize its tests when a deployment loop searches, scores, and selects patches.
  2. 23 min Episode 2Gaming the Verifier: Test-Suite, Judge, and KPI-Proxy Capture in Agentic Tool-UseTrace one duplicate-charge support case through three unchanged graders and identify the independent evidence each verdict leaves out.
  3. 24 min Episode 3RLAIF and Self-Grading Loops: When the Generator and the Grader Share a MindA report writer and its critic can agree for the same wrong reason; independent evidence makes self-review useful without mistaking agreement for truth.
  4. 23 min Episode 4Reward Tampering and Wireheading: When the Agent Can Reach the Reward ChannelFollow a coding assistant's evaluation channel from task output to accepted score, and place integrity controls beyond the assistant's authority.
  5. 24 min Episode 5Detecting and Containing Reward Hacking in Production: The Divergence PlaybookTurn apparently excellent scores into a controlled release decision using independent outcome evidence, search history, and reward-channel integrity.