Research

Watch · narrated whiteboard episodesL3

Adversarial Inputs and Jailbreaks: The Optimization View

A jailbreak is not a clever sentence but the solution to a discrete optimization problem over tokens. This threat lab formalizes the adversarial objective, derives gradient-guided coordinate search, explains transferability and universality as shared representation geometry, maps the jailbreak families to the single mechanism each exploits, and bounds what robustness can promise. Grounded in the primary adversarial-ML literature, each primitive paired with its defense.

Murali Chillakuru·5 episodes
  1. 14 min Episode 1The Adversarial Objective: Formalizing a Jailbreak as a Loss over the InputA moderator and an expert reframe jailbreaks from clever phrases into the minimizers of a loss function — showing that once you see the math, the defenses follow directly from the structure of the attack.
  2. 12 min Episode 2Greedy Coordinate Search: Why Token-Level Search Beats Naive Gradient DescentA moderator and an expert dissect the Greedy Coordinate Gradient algorithm — showing how gradient nomination and greedy search combine to find adversarial suffixes that reliably jailbreak language models.
  3. 11 min Episode 3Universality and Transfer: Shared Representation Geometry as the Reason Suffixes MoveA moderator and an expert examine why a single adversarial suffix can jailbreak many different prompts and models it never saw during optimization — and what the underlying geometry explains about both the threat and the defense.
  4. 11 min Episode 4A Taxonomy of Jailbreak Families: One Mechanism Behind Obfuscation, Role-Play, and Many-ShotA moderator and an expert show that the dozens of named jailbreak styles — role-play, obfuscation, cipher, many-shot — are all surface variations on a single move: pushing the input into the region of the model's input space where refusal probability is low.
  5. 13 min Episode 5Robustness and Its Limits: Adversarial Training, Classifiers, and Why Provable Bounds Stay SmallA moderator and an expert examine every practical jailbreak defense, show that each one raises the attacker's cost without closing the door, and build the honest case for defense in depth as the only workable posture.