Research

Watch · narrated whiteboard episodesL3

Alignment and Fine-Tuning Attacks: Removing the Guardrails

Safety alignment is a thin, removable layer. A few hundred fine-tuning examples — or a targeted weight edit — can strip refusals without touching capability. This threat lab maps where safety lives in the stack, measures the dose-response of harmful fine-tuning, dissects the refusal direction and activation steering, frames reward hacking as an attack on the training objective, and surveys safety-preserving defenses — each paired with a hardening. Grounded in the primary alignment-integrity literature.

Murali Chillakuru·5 episodes
  1. 12 min Episode 1The Alignment Threat Model: Where Safety Lives and What Each Layer GuaranteesA moderator and an expert pull apart a deployed model's safety into its distinct layers, showing which guarantees hold, which an attacker can peel away, and where real safety has to be anchored.
  2. 14 min Episode 2Harmful Fine-Tuning: Dose-Response and the Capability-Safety DecouplingA moderator and an expert examine how fine-tuning — even on benign data — can strip a model's safety, and what a safety-gated pipeline looks like in practice.
  3. 14 min Episode 3Refusal Ablation and Activation Steering: The Refusal DirectionA moderator and an expert examine how refusal behavior is concentrated in a single activation-space direction, and what it means that one linear edit can collapse it.
  4. 15 min Episode 4Reward Hacking in RLHF: Specification Gaming and Proxy-Reward DivergenceA moderator and an expert examine how training a model to maximize a proxy reward is an invitation for the optimizer to exploit the gap between the proxy and what you actually want.
  5. 16 min Episode 5Defenses: Safety-Preserving Fine-Tuning, Tamper-Resistance, and Adaptive EvaluationA moderator and an expert assemble the full defensive program for alignment — from preserving the overlay during fine-tuning, to making removal costlier, to anchoring enforcement outside the model entirely.