Reasoning redundancy can now be scored step by step
Reasoning redundancy can now be scored step by step, and picking fine-tuning data by that score saves inference; the evaluation is the authors’ own.
ImportânciaLocalEvidênciaE2 não replicadaTratamentoRápido
A reasoning chain can now be scored step by step for wasted work, and picking fine-tuning data by that score makes inference cheaper with almost no loss in task accuracy.
Models that re-derive steps they have already settled burn compute and add latency; until now the only checks were reading trajectories by hand or comparing total lengths, neither of which says which step was the redundant one.
On the redundancy class of the PRMBench dataset, the score identifies redundant steps more than 10 points more accurately than embedding-similarity and information-gain baselines, and it tracks the actual reasoning length of QwQ-32B, DeepSeek-R1-Distill-Qwen-32B and GPT-4.1. It comes from the SLIDER framework, which uses partial information decomposition to split what two adjacent reasoning steps contribute to the final answer into unique, redundant and synergistic parts, and flags a step as repetitive when redundancy dominates.
The evaluation is the authors’ own and covers that redundancy dataset; the paper was submitted on September 30 and accepted to an ICLR 2026 workshop on logical reasoning in large models.