Watch · narrated walkthroughs
AI-output detection and watermarking are security controls, and like all controls they have an attacker. This threat lab analyzes the watermark/detector game formally: the detection threat model, how green-list text watermarks work and their robustness/quality trade-off, paraphrase and spoofing evasion, cryptographic content provenance versus statistical watermarks, and the impossibility-leaning limits of detection — each paired with a hardening. Grounded in the primary watermarking and detectability literature.