Research

Watch · narrated whiteboard episodesL3

Watermarking and Provenance Evasion: The Attribution Arms Race

AI-output detection and watermarking are security controls, and like all controls they have an attacker. This threat lab analyzes the watermark/detector game formally: the detection threat model, how green-list text watermarks work and their robustness/quality trade-off, paraphrase and spoofing evasion, cryptographic content provenance versus statistical watermarks, and the impossibility-leaning limits of detection — each paired with a hardening. Grounded in the primary watermarking and detectability literature.

Murali Chillakuru·5 episodes
  1. 16 min Episode 1The Detection Threat Model: Evasion, Spoofing, and Scrubbing GoalsA moderator and a security expert lay the groundwork for AI-text detection — what a detector actually promises, the three things an adversary wants, and why the whole problem is an adversarial game where false accusations may be the worst outcome of all.
  2. 16 min Episode 2How Text Watermarks Work: Green-List Schemes and the Robustness-Quality Trade-offA moderator and a security expert open up the machinery of statistical text watermarking — the green-list trick that biases word choice, the statistical test that reads it back, and the unavoidable trade-off between a robust mark and natural writing.
  3. 17 min Episode 3Evasion and Spoofing: Paraphrase Attacks, Watermark Stealing, and ForgeryA moderator and a security expert walk through the two families of attack on text watermarks — evasion that scrubs a real mark away, and spoofing that forges a mark onto innocent text — and explain why the second is the graver danger.
  4. 16 min Episode 4Content Provenance: Cryptographic Signing Versus Statistical WatermarksA moderator and a security expert contrast watermarking with cryptographic content provenance — signed records of origin like C2PA — showing that they answer two different questions, break in different ways, and are strongest when composed together.
  5. 18 min Episode 5What Detection Can and Cannot Promise: Impossibility-Leaning ResultsA moderator and a security expert close the series with the hard theory — why detection is bounded by how close human and machine text can get, the false-positive/false-negative bind, impossibility-style results, and how to build trustworthy systems that respect the limits.