Alignment is purpose-agnostic — and works as a censor's toolkit
Alignment methods make a model serve any principal's will without guaranteeing good intent; the authors back transparency, verifiable alignment, and model pluralism.
BedeutungLokalBeweisE2 nicht repliziertAufbereitungSchnell
The same alignment techniques built to make models safe can just as easily censor or distort information: alignment methods are purpose-agnostic, making a model serve someone's will with nothing in the methodology guaranteeing good intent.
Alignment has largely been assumed to be a safety measure; Sarah Ball and Phil Hackemann sort control into three levers — pretraining data filtering, post-training alignment, and inference-time intervention — with cost falling and ease of change rising down the stack, making lower layers easier to abuse unilaterally.
This is not hypothetical: China's cyberspace regulator requires providers to maintain refusal datasets, and Elon Musk publicly said he would "fix" Grok outputs he disagreed with, with reported behavior shifts traced to system-prompt changes. The analysis won an Outstanding Position Paper Award at ICML 2026; the authors do not call for stopping alignment, but back transparency, verifiable alignment, and model pluralism.
This is a position paper's argument plus public examples — author-reported views, with no third-party replication of the quantified risk.