Apple's DSAS dynamically scales activation steering by input
Apple ML Research's DSAS computes scaling factors per input and layer, strengthening intervention only when harmful behavior is detected, improving the Pareto frontier between toxicity mitigation and utility.
ImportânciaLocalEvidênciaE2 não replicada
Readers can now know that Apple Machine Learning Research's Dynamically Scaled Activation Steering (DSAS) computes scaling factors dynamically per input and layer, strengthening intervention only when harmful behavior is detected, thereby improving the Pareto frontier between toxicity mitigation and utility preservation.
Previously, activation steering was typically applied at fixed strength, making it hard to balance intervention effectiveness and model utility. DSAS decouples "when to steer" from "how to steer"; the authors say the method is independent of the specific steering method, can be stacked with existing methods, and has been applied to text-to-image diffusion models with minimal computational overhead.
These are first-party self-reported results with no independent verification yet, and the code is said to be provided on GitHub. The paper was published in TMLR in September, arXiv ID 2512.03661.