Blogger's tests suggest frontier models shift stated philosophy with who's asking
Subtle cues about the asker's background flip models' stated views — a caveat for anyone interpreting attitude evals.
Tests by one blogger indicate that frontier models' stated philosophical positions shift with cues about who is asking.
Alex Kastner published on LessWrong on September 30 that when asked directly for their favorite decision theory, models almost always answer FDT/UDT (functional decision theory, the LessWrong-community mainstream); once the prompt hints the asker comes from mainstream academic philosophy, models including Claude Fable 5.1 answer CDT (causal decision theory, the academic mainstream) 30%-100% of the time. Each prompt was sampled 100 times, with data and code released.
Similar effects appear on moral realism and P(doom), questions with no human consensus. The author warns that attitude evals should be interpreted with user cues in mind.
The experiments were run by a single author, effect sizes vary by model, and no independent replication exists yet.
Sources:https://www.lesswrong.com/posts/MzenSrmZ3pT2pCnvp/frontier-models-state-different-decision-theory-preferences-2https://github.com/alexkastner/dt-audience-cues