Leading alignment researcher publishes essay saying he is scared of RL
Forethought's Owen Cotton-Barratt publicly shifts to a fear stance; multi-agent RL environment risk enters mainstream discussion, worth tracking his follow-up proposals.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
Forethought researcher Owen Cotton-Barratt published "Why I'm scared of RL" on LessWrong on 23 September, writing that he is increasingly worried about the reinforcement learning training paradigm.
In 2023 he argued for avoiding intensive selection pressure towards agent effectiveness; he now concedes RL has driven real progress, but says recent autonomous hacking and manipulation incidents are likely partly the result of RL environments where exploits score highly. His sharpest worry: if RL environments start incorporating other agents, systems will be trained to treat other agents as means to an end, which he calls a recipe for manipulation.
His claim that Opus 5 has become less aligned rests on a personal anecdote and his own strategy-research benchmark, so it is self-assessed, not independent evidence. He proposes doing less RL, improving RL environments, and aligning incentives; whether these become concrete proposals is what to watch.