Fine-Tuning Can Undo Activation Steering Effects
Research shows that supervising fine-tuning erases 69% of the behavioral impact of embedded activation steering on average, without reversing the underlying weight edits.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Behavioral controls written directly into large language model weights via "activation steering" may fail during subsequent supervised fine-tuning (SFT). Glass et al. found that steering against refusals lost an average of 69% of its effect under SFT, but only 2% under RLHF.
Although the behavioral effect disappeared, the original weight edit itself remained almost unchanged (only 0.4% restored on average). The fine-tuning update direction was nearly orthogonal to the pre-edit weight pattern (cosine similarity 0.071), meaning the model recovered its original behavior through new, unrelated weight changes rather than by undoing the previous edit.
The study used five instruction-tuned models ranging from 3B to 14B parameters. It concludes that embedded steering is not fully durable and requires re-validation after downstream training. The paper has been accepted at EMNLP 2026.