Research

Watch · narrated whiteboard episodesL3

The Reasoning Trace as an Attack Surface

Extended-reasoning models think out loud, and it is tempting to trust that visible chain-of-thought as the real reason for an answer. This threat lab shows why that trust is unearned: the trace can be an unfaithful rationalization, an injection target, an exfiltration channel, and a timing side channel — all at once. Each exposure mode is taught by the assumption it breaks, the mechanism that makes it work, and the assumption-free control that does not trust the trace. Product-agnostic and grounded in the primary faithfulness and monitorability literature, tied back to the AI-agent stack every time.

Murali Chillakuru·5 episodes
  1. 19 min Episode 1What a Reasoning Trace Is and Is Not: Faithfulness, Post-Hoc Rationalization, and the Monitorability AssumptionA stage walkthrough of why a model's reasoning trace is a generated artifact, not a window into computation, and why trusting it as an oversight signal is itself an attack surface.
  2. 18 min Episode 2Injecting the Scratchpad: When Retrieved or Tool-Returned Content Steers the Reasoning, Not Just the AnswerA stage walkthrough of how untrusted documents can bend a model's intermediate reasoning — and launder the steer through a clean-looking trace that oversight then approves.
  3. 17 min Episode 3The Trace as Exfiltration Channel: Encoded Reasoning and Steganographic ScratchpadsA stage walkthrough of how a consumed reasoning trace becomes an outbound channel — carrying secrets and hidden computation in its surface form — and why paraphrasing before consumption is the strongest known control.
  4. 16 min Episode 4Reasoning-Budget Side Channels: Token-Count and Latency as an Observable OracleA stage walkthrough of why hiding a model's chain-of-thought doesn't hide how hard it thought — and how token counts and latency become an oracle unless effort is padded and quantized.
  5. 16 min Episode 5Defenses That Don't Trust the Trace: Faithfulness Testing, Trace Isolation, and Monitor LimitsThe synthesizing finale: an architecture that stays safe whether or not any given reasoning trace is honest — measure faithfulness, isolate the trace, bound the monitor, backstop the action, close the non-text channels, and measure forever.