Watch · narrated whiteboard episodesL3
Extended-reasoning models think out loud, and it is tempting to trust that visible chain-of-thought as the real reason for an answer. This threat lab shows why that trust is unearned: the trace can be an unfaithful rationalization, an injection target, an exfiltration channel, and a timing side channel — all at once. Each exposure mode is taught by the assumption it breaks, the mechanism that makes it work, and the assumption-free control that does not trust the trace. Product-agnostic and grounded in the primary faithfulness and monitorability literature, tied back to the AI-agent stack every time.