Training-free boost lifts dLLM reasoning by 16%
New framework RFG enables diffusion LLMs to self-improve at inference time, boosting performance by up to 16.1% without additional training or data.
ImportanceMaterialEvidenceE2 unreplicatedWrite-upQuick
Diffusion Large Language Models (dLLMs) can now self-improve during inference using the new Reward-Free Guidance (RFG) framework, achieving up to 16.1% performance gains.
Previously, enhancing dLLM capabilities required costly post-training with extra data and supervision. RFG introduces a training-free method that derives guidance signals directly from model checkpoints, addressing the lack of well-defined signals for partially masked intermediate states.
The study, led by Stanford researchers, theoretically demonstrates that reward signals can be parameterized via log-likelihood ratios between policy and reference models. Experiments show these gains rival or surpass resource-intensive reinforcement learning techniques despite requiring no training.
This is a preprint result; independent reproduction is pending.