Research

Watch · narrated whiteboard episodesL3

Goal-Drift Detection in Autonomous Agents

An autonomous agent that quietly stops pursuing the goal it was given — while still looking busy and authorized — is one of the hardest failures to catch. This series builds runtime goal-drift detection from the ground up: what drift is and why non-determinism makes it inevitable, how to specify the intended behavior and reference policy to drift from, how to detect divergence at runtime with trajectory distance, reward-model monitors, and statistical change detection, how to survive the base-rate problem with drift budgets and calibrated escalation, and how to close the loop with correction, rollback, and human re-grounding without halting the fleet. Grounded in NIST AI RMF, OWASP LLM Top 10, the OWASP Agentic Security Initiative, MITRE ATLAS, the AI-safety literature on reward hacking, and classical statistical change-point detection.

Murali Chillakuru·5 episodes
  1. 10 min Episode 1The Goal-Drift Problem: What Drift Is, Why Non-Determinism Makes It Inevitable, and the Detection ObjectiveA stage conversation on why autonomous agents inevitably wander from their goals, why static controls can't see it, and how to frame drift detection as a calibrated change-detection problem.
  2. 9 min Episode 2Specifying Intended Behavior: Goal Representations, Invariants, and the Reference Policy to Drift FromA stage conversation on building the reference a drift detector compares against — goal representations, hard invariants, and the specification gap every reference inherits.
  3. 9 min Episode 3Detecting Drift at Runtime: Trajectory Distance, Reward-Model Monitors, and Statistical Change DetectionA stage conversation on converting a reference into live alarms — measuring departure, scoring conformance, and deciding when a wobble has become a persistent shift.
  4. 10 min Episode 4The Base-Rate Problem: False-Positive Cost, Drift Budgets, and Calibrated Escalation ThresholdsA stage conversation on why an accurate drift detector can still flood you with false alarms, and how budgets, base-rate calibration, and precision keep it usable.
  5. 9 min Episode 5Closing the Loop: Correction, Rollback, and Human Re-Grounding Without Halting the FleetA stage conversation on turning a drift alarm into a targeted correction — a graduated ladder, a stable control loop, and containment that never stops the whole fleet.