Key-position confound inflates no-CoT reasoning depth on benchmark
Moving the initial state to the prompt's end cuts measured reasoning depth by a median 16%, a caveat for reading looped-transformer advantages.
Moving the initial state to the end of the prompt cuts GPT-6.1 Sol's no-chain-of-thought reasoning depth by a median 16%.
Neel Nanda's no-cot-bench had shown Astra, suspected to be a looped transformer, solving problems at roughly 8.6x the odds of Fable 5.1. Two researchers identify a confound: the key usually sits at the start of the prompt, so a model can compute while reading, meaning the benchmark measures both serial depth and the ability to spread computation across tokens.
They built key-last versions of six serial state-tracking tasks, with an unrelated-question control to rule out parsing difficulty. The median depth loss across 12 tasks was 16%, ranging from 9% to 45%. The authors tested only GPT-6.1 Sol, saying they lacked the compute to rerun other models.
Sources:https://www.lesswrong.com/posts/ehF4kej38fkhb5cT3/a-key-position-confound-in-no-cot-bench-and-implications-forhttps://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-do-a-concerning-amount-with-no-chain-of-thought