RL-trained LoRA rewriting policy aligns SFT data distribution, non-downstream degradation drops across three backbones
Authors self-report: a lightweight LoRA rewriting policy trained with reinforcement learning, under a task-consistency hard gate, optimizes QA-style distribution alignment and semantic diversity, with upstream-downstream gains roughly matching standard SFT across three instruction-tuned backbones.
ImportanceLocalEvidenceE2 unreplicated
A lightweight LoRA rewriting policy trained with reinforcement learning can rewrite SFT data, optimizing QA-style distribution alignment and semantic diversity under a task-consistency hard gate; the authors self-report that across three instruction-tuned backbones, upstream-downstream gains are roughly on par with standard SFT, and non-downstream benchmark degradation is reduced in all evaluation settings.
Previously, SFT data rewriting lacked this kind of formalization, making it hard to control rewriting quality and downstream cost at the same time.
The result is self-reported by the authors: on three instruction-tuned backbones, upstream-downstream gains are roughly on par with standard SFT, and non-downstream benchmark degradation is reduced in all evaluation settings; cross-domain reuse of the rewriting policy is only preliminary evidence.
Boundary: the results have not been independently reproduced; the work is the arXiv preprint 2602.11220, first submitted February 11 and updated to v2 on September 16. Source: arXiv abstract page ↗