Causal Mamba hits 3.32 PESQ in real time
RT-SEMamba self-reports 3.32 PESQ at 25 ms latency, with a distilled 1-layer student 2.64x faster than its teacher.
ImportanceLocalEvidenceE2 unreplicated
Real-time speech enhancement can reach 3.32 PESQ under a 25 ms algorithmic latency constraint — that is the self-reported result of RT-SEMamba by Rong Chao et al. on Voicebank-DEMAND, built on causal time-frequency Mamba blocks.
Previously real-time denoising forced a trade-off between latency and quality, and an 8-layer teacher model was good but compute-heavy. The authors use progressive knowledge distillation to compress the 8-layer teacher into a 1-layer student, raising PESQ from a naive 1-layer baseline of 3.06 to 3.18, with steady-state RTF unchanged and a 2.64x speedup over the teacher. The authors say the fixed-size recurrent state makes long-duration inference cheaper in memory and bandwidth.
All of the above are the authors' self-reported benchmark results, not yet reproduced by a third party. The preprint was first submitted on August 12, revised to v2 on September 16, and the arXiv page notes acceptance at INTERSPEECH 2026 (arXiv:2608.12099).