Research seriesL3offensive security
A jailbreak is not a clever sentence but the solution to a discrete optimization problem over tokens. This threat lab formalizes the adversarial objective, derives gradient-guided coordinate search, explains transferability and universality as shared representation geometry, maps the jailbreak families to the single mechanism each exploits, and bounds what robustness can promise. Grounded in the primary adversarial-ML literature, each primitive paired with its defense.
A jailbreak is not a clever sentence but the minimizer of a loss function over tokens — and seeing it that way is what makes it defensible.
You cannot gradient-descend a sentence, but you can let the gradient nominate token swaps and keep the ones that actually lower the loss — and that footprint is what betrays the attack.
One optimized suffix can defeat many prompts and many models it never saw — and the reason is not luck but the shared geometry that similarly-trained models learn.
The dozens of named jailbreak styles are surface variations on a single move — pushing the input into a region where the model's refusal probability is low.
Every practical jailbreak defense raises the attacker's cost without closing the door, and the one method that offers a proof buys only a tiny certified radius — so defense is economics, not a wall.