Activation steering largely fails on instruction-tuned models, at a fluency cost
Teams relying on activation steering to control LLM outputs should note: Apple researchers report sharply reduced effectiveness on instruction-tuned models, at a fluency cost.
ImportanceLocalPreuvesE2 non réplicableTraitementRapide
Activation steering methods are far less effective on instruction-tuned models than on their base counterparts, and often impose a steep fluency cost — a finding from researchers at Apple Machine Learning and Pompeu Fabra University.
Teams that previously relied on activation steering to control LLM outputs — such as removing toxic concepts — may have overestimated how well these weight-free, internal-activation interventions transfer to instruction-tuned models; meanwhile, prompting and full supervised fine-tuning work for concept injection but handle concept removal poorly.
The result is the authors' own evaluation; the paper was published as a workshop paper at BlackBoxNLP 2026, the arXiv version was first submitted on June 10, and Apple's research page lists it as published in September 2026.
Boundary: the abstract does not list the specific models or sample sizes, so extrapolation calls for caution, and no third-party replication exists yet.