TopicSpectrum AI ResearchMurali Chillakuru
/
About
All Research tracks

Defensive AI Security

Feature dictionaries, sparse autoencoders, and activation probes are making model internals partially legible at runtime — and that raises a defensive question: can reading the internals detect deception, jailbreak states, or unsafe behavior before it reaches the output? This paper treats interpretability as a monitoring control rather than a research curiosity, and asks the harder question every detector must answer: what are its false-negative limits under an adaptive adversary? The organizing contribution is a Probe Assurance Ledger that separates 'evidence' from 'guarantee' for each internal-signal detector. Product-agnostic, grounded in the primary probing and feature-learning literature, and tied back to the AI-agent stack every time.

1 series · 5 articles

Interpretability for Security