Award-winning paper warns alignment techniques are becoming a censor's toolkit
An ICML award-winning position paper argues alignment methods are purpose-agnostic and can serve censorship by states or vendors; the authors back transparency and model pluralism.
Original event 2026-09-28
A position paper that won an Outstanding Position Paper Award at ICML 2026 warns that the same alignment techniques built to make models safe can just as easily be used to censor or distort information.
Authors Sarah Ball and Phil Hackemann argue alignment methods are purpose-agnostic: they make a model serve someone's will, with nothing in the methodology guaranteeing good intent. They sort control into three levers — pretraining data filtering, post-training alignment, and inference-time intervention — with cost falling and ease of change rising down the stack.
The paper says this is not hypothetical: China's cyberspace regulator requires providers to maintain refusal datasets, and Elon Musk publicly said he would "fix" Grok outputs he disagreed with, with reported behavior shifts traced to system-prompt changes. The authors do not call for stopping alignment; they back transparency, verifiable alignment, and model pluralism.