Watch · narrated walkthroughs
A jailbreak is not a clever sentence but the solution to a discrete optimization problem over tokens. This threat lab formalizes the adversarial objective, derives gradient-guided coordinate search, explains transferability and universality as shared representation geometry, maps the jailbreak families to the single mechanism each exploits, and bounds what robustness can promise. Grounded in the primary adversarial-ML literature, each primitive paired with its defense.