Research

Watch · narrated walkthroughs

Alignment and Fine-Tuning Attacks: Removing the Guardrails

Safety alignment is a thin, removable layer. A few hundred fine-tuning examples — or a targeted weight edit — can strip refusals without touching capability. This threat lab maps where safety lives in the stack, measures the dose-response of harmful fine-tuning, dissects the refusal direction and activation steering, frames reward hacking as an attack on the training objective, and surveys safety-preserving defenses — each paired with a hardening. Grounded in the primary alignment-integrity literature.

Murali Chillakuru·5 episodes
  1. 1
  2. 2
  3. 3
  4. 4
  5. 5