Research

Research seriesL3offensive security

Watermarking and Provenance Evasion: The Attribution Arms Race

AI-output detection and watermarking are security controls, and like all controls they have an attacker. This threat lab analyzes the watermark/detector game formally: the detection threat model, how green-list text watermarks work and their robustness/quality trade-off, paraphrase and spoofing evasion, cryptographic content provenance versus statistical watermarks, and the impossibility-leaning limits of detection — each paired with a hardening. Grounded in the primary watermarking and detectability literature.

Murali Chillakuru·5 articles
  1. 1
    The Detection Threat Model: Evasion, Spoofing, and Scrubbing Goals

    AI-content detection is a security control, and like every control it has a motivated adversary — so its guarantees must be stated against attackers, not against honest users.

  2. 2
    How Text Watermarks Work: Green-List Schemes and the Robustness-Quality Trade-off

    A text watermark is a statistical thumb on the scale — bias which tokens the model prefers, then test for that bias later — and its whole design lives on a robustness-versus-quality curve.

  3. 3
    Evasion and Spoofing: Paraphrase Attacks, Watermark Stealing, and Forgery

    A statistical watermark can be washed out by rewriting and, more dangerously, forged onto an innocent author — so its adversary is not only the evader but the framer.

  4. 4
    Content Provenance: Cryptographic Signing Versus Statistical Watermarks

    Statistical watermarks ask 'does this look AI-made?'; cryptographic provenance asks 'who made this, and has it changed?' — different questions with very different security.

  5. 5
    What Detection Can and Cannot Promise: Impossibility-Leaning Results

    As generators improve, the statistical gap detectors rely on shrinks toward zero — turning reliable detection from an engineering problem into a provably hard one.