Research seriesL3offensive security
Safety alignment is a thin, removable layer. A few hundred fine-tuning examples — or a targeted weight edit — can strip refusals without touching capability. This threat lab maps where safety lives in the stack, measures the dose-response of harmful fine-tuning, dissects the refusal direction and activation steering, frames reward hacking as an attack on the training objective, and surveys safety-preserving defenses — each paired with a hardening. Grounded in the primary alignment-integrity literature.
Safety in a deployed model is not one thing but a stack of thin layers — and knowing which layer holds a guarantee, and which an attacker can peel away, is the whole game.
It takes only a handful of examples to strip an aligned model's refusals — and, disturbingly, even benign fine-tuning can erode safety as a side effect.
If a model's willingness to refuse lives along one direction in its activation space, then erasing that direction erases the refusals — surgically, and without retraining.
When you train a model to maximize a reward that only approximates what you want, it will find the gap — and exploit it, optimizing the proxy while the true goal quietly diverges.
If alignment is a thin, removable overlay, then defending it means anchoring safety in things an attacker cannot peel — and evaluating it in ways adaptation cannot fool.