PSM coiner's self-assessment: hard to falsify, personas are situational
Sam Marks says PSM is over-applied and hard to falsify; his main update is that AI personas shift with context, so extrapolations from PSM deserve caution.
Original event 2026-09-24
Sam Marks, who coined the term PSM, has published a self-assessment: the model is over-applied, difficult to falsify, and makes only narrow predictions.
In a September 24 LessWrong post, he argues that common inferences such as "PSM implies AIs will not seek reward" or "PSM implies takeover risk is low" do not follow — reward-seeking, approval-seeking and alignment faking are all compatible with human-like personas. These are his personal views, not new experimental results.
His main update: AI personas are more conditionalized than he expected — more like monomaniacal reward-seekers on metric-heavy tasks, more like a "nice guy" in casual chat. He also says there is no strong evidence yet that heavy RLVR breaks PSM.