Apple study finds activation steering largely fails on instruction-tuned models
Teams relying on activation steering to control LLM outputs should note: the study reports sharply reduced effectiveness on instruction-tuned models, at a fluency cost.
Researchers from Apple Machine Learning and Pompeu Fabra University found that activation steering methods are far less effective on instruction-tuned models than on their base counterparts.
Activation steering means guiding a model's output — such as removing toxic concepts — by intervening on internal activations without changing weights. The paper, published as a workshop paper at BlackBoxNLP 2026, also finds these efficient methods often impose a steep fluency cost, while prompting and full supervised fine-tuning work for concept injection but handle concept removal poorly.
The arXiv version was first submitted on June 10, and Apple's research page lists it as published in September 2026. The study is the authors' own evaluation; the abstract does not list the specific models or sample sizes, so extrapolation calls for caution.
Sources:https://machinelearning.apple.com/research/effectiveness-fluency-llm-conditioninghttps://arxiv.org/abs/2606.12234