Research

Watch · narrated whiteboard episodesL3

Mechanistic Interpretability as a Security Control

Feature dictionaries, sparse autoencoders, and activation probes are making model internals partially legible at runtime — and that raises a defensive question: can reading the internals detect deception, jailbreak states, or unsafe behavior before it reaches the output? This paper treats interpretability as a monitoring control rather than a research curiosity, and asks the harder question every detector must answer: what are its false-negative limits under an adaptive adversary? The organizing contribution is a Probe Assurance Ledger that separates 'evidence' from 'guarantee' for each internal-signal detector. Product-agnostic, grounded in the primary probing and feature-learning literature, and tied back to the AI-agent stack every time.

Murali Chillakuru·5 episodes
  1. 17 min Episode 1Reading the Machine: Features, Sparse Autoencoders, and What Interpretable Actually Buys a DefenderA moderator and a staff-level researcher take a defensive lens on model interpretability — what reading internals buys a security engineer, and why every internal signal is evidence with an error rate, never a guarantee.
  2. 16 min Episode 2Probes as Detectors: Classifying Jailbreak, Refusal, and Unsafe States From ActivationsA moderator and a staff-level researcher turn an activation probe into a deployable security detector — the states worth reading, how to validate like a data scientist, and the base-rate trap that sinks most probes in production.
  3. 16 min Episode 3Detecting Deception From the Inside: Promise, Evidence, and the Adaptive-Adversary ProblemA moderator and a staff-level researcher separate three things that get blurred — the promise that a lie leaves an internal signature, the evidence that actually exists, and the adaptive-adversary problem that makes deception detection categorically harder than reading an externally-induced state.
  4. 15 min Episode 4Runtime Activation Monitoring: Latency, Coverage, and the Cost of Watching InternalsA moderator and a staff-level researcher treat internal monitoring as a systems problem — where the latency goes, inline versus shadow, and a cost-coverage frontier that says you buy the most safety by concentrating a limited budget, not by watching everything.
  5. 23 min Episode 5Interpretability's Limits as a Control: False Negatives, Goodharting the Probe, and Layered DefenseAn internal alarm can stop a risky action, but a quiet detector cannot authorize it: a worked assurance case for a refund assistant.