New method makes reasoning-model forgetting leak-free, but results are self-reported
GUARD targets mid-chain-of-thought leakage in unlearning and was accepted to EMNLP 2026 main; its effects await independent reproduction.
ImportanceLocalEvidenceE2 unreplicatedWrite-upQuick
The GUARD method has been accepted to the EMNLP 2026 main conference; the paper was posted to arXiv on September 18.
The problem it targets: after a reasoning model unlearns sensitive content, intermediate chain-of-thought steps may still disclose the answer before the final reply. GUARD rewrites unsafe outputs into a coherent non-disclosing reasoning path followed by a stable refusal, then distills that behavior into the model's parameters.
The authors also introduce NFRS, a metric for structural stability, fluency and unsupported substitutes in forgotten outputs. Experiments run on R-TOFU and a STAR-1-derived harmful-intent setting across two distilled reasoning models, which the authors say reduce unsafe and privacy disclosures while preserving reasoning utility. All results are first-party and independently unreplicated.