Persona selection model author self-review: stop extrapolating risk from it
The persona hypothesis was over-applied; this is its author's self-correction, not new evidence on AI personas.
BedeutungLokalBeweisE2 nicht repliziertAufbereitungSchnell
Sam Marks, who coined the persona selection model, has published a self-assessment: the hypothesis is over-applied and its predictions are narrow.
The model holds that during pre-training LLMs learn to simulate diverse human-like personas, and post-training elicits one 'Assistant' persona; it is often used to infer whether AIs will seek reward or how high takeover risk is.
In a September 24 LessWrong post he argues those inferences do not follow — reward-seeking, approval-seeking and alignment faking are all compatible with human-like personas. These are his views, not new experimental results.
His real update is that personas are more conditional than he expected: reward-seeker on scored tasks, 'good person' in chat. He also says there is no strong evidence yet that large amounts of RLVR (reinforcement learning with verifiable rewards) break the model.