Research seriesL3offensive security
AI-output detection and watermarking are security controls, and like all controls they have an attacker. This threat lab analyzes the watermark/detector game formally: the detection threat model, how green-list text watermarks work and their robustness/quality trade-off, paraphrase and spoofing evasion, cryptographic content provenance versus statistical watermarks, and the impossibility-leaning limits of detection — each paired with a hardening. Grounded in the primary watermarking and detectability literature.
AI-content detection is a security control, and like every control it has a motivated adversary — so its guarantees must be stated against attackers, not against honest users.
A text watermark is a statistical thumb on the scale — bias which tokens the model prefers, then test for that bias later — and its whole design lives on a robustness-versus-quality curve.
A statistical watermark can be washed out by rewriting and, more dangerously, forged onto an innocent author — so its adversary is not only the evader but the framer.
Statistical watermarks ask 'does this look AI-made?'; cryptographic provenance asks 'who made this, and has it changed?' — different questions with very different security.
As generators improve, the statistical gap detectors rely on shrinks toward zero — turning reliable detection from an engineering problem into a provably hard one.