New metric quantifies reasoning redundancy, beating baselines by over 10 points
A paper proposes Step-RRI to quantify redundant reasoning and guide fine-tuning data selection; results are author-run.
A paper submitted on September 30 proposes Step-RRI, an index that on the redundancy class of the PRMBench dataset improves step-level redundancy identification accuracy by more than 10 points over embedding-similarity and information-gain baselines.
The metric comes from the SLIDER framework: it uses Partial Information Decomposition, an information-theory method, to split the information two consecutive reasoning steps carry about the final answer into unique, redundant and synergistic components, flagging a step as repetitive when redundancy dominates. The paper reports that a trajectory-level version strongly correlates with actual reasoning length on QwQ-32B, DeepSeek-R1-Distill-Qwen-32B and GPT-4.1.
The authors also show that selecting fine-tuning data by this index can improve reasoning efficiency while largely preserving task performance. The paper is accepted at the ICLR 2026 workshop on logical reasoning of large language models; the evaluation is the authors' own and covers only that redundancy dataset.
Sources:https://arxiv.org/abs/2610.00571