Papers › Robust One-step Speech Enhancement via Consistency Distillation

Robust One-step Speech Enhancement via Consistency Distillation

8 Jul 2025arXiv:2507.05688archive 2025-07-28

Liang Xu, Longfei Felix Yan, W. Bastiaan Kleijn

Diffusion models have shown strong performance in speech enhancement, but their real-time applicability has been limited by multi-step iterative sampling. Consistency distillation has recently emerged as a promising alternative by distilling a one-step consistency model from a multi-step diffusion-based teacher model. However, distilled consistency models are inherently biased towards the sampling trajectory of the teacher model, making them less robust to noise and prone to inheriting inaccuracies from the teacher model. To address this limitation, we propose ROSE-CD: Robust One-step Speech Enhancement via Consistency Distillation, a novel approach for distilling a one-step consistency model. Specifically, we introduce a randomized learning trajectory to improve the model's robustness to noise. Furthermore, we jointly optimize the one-step model with two time-domain auxiliary losses, enabling it to recover from teacher-induced errors and surpass the teacher model in overall performance. This is the first pure one-step consistency distillation model for diffusion-based speech enhancement, achieving 54 times faster inference speed and superior performance compared to its 30-step teacher model. Experiments on the VoiceBank-DEMAND dataset demonstrate that the proposed model achieves state-of-the-art performance in terms of speech quality. Moreover, its generalization ability is validated on both an out-of-domain dataset and real-world noisy recordings.

PaperPDFConference PDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Speech Enhancement

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Enhancement VoiceBank + DEMAND ROSE-CD(PESQ) CBAK 3.37 #1 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD(PESQ) COVL 4.30 #1 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD(PESQ) CSIG 4.63 #1 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD(PESQ) ESTOI 0.83 #1 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD(PESQ) PESQ (wb) 3.99 #1 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD(PESQ) Para. (M) 65 #1 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD(PESQ) SI-SDR 0.40 #1 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD(PESQ) SSNR 0.927 #1 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD(PESQ) STOI 92.6 #1 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD CBAK 3.33 #13 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD COVL 4.04 #13 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD CSIG 4.523 #13 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD ESTOI 0.87 #13 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD PESQ (wb) 3.49 #13 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD Para. (M) 65 #13 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD SI-SDR 17.80 #13 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD SSNR 3.34 #13 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND ROSE-CD STOI 94.73 #13 of 42 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND ROSE-CD DNSMOS 3.48 #1 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND ROSE-CD DNSMOS BAK 4.34 #1 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND ROSE-CD DNSMOS OVRL 3.70 #1 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND ROSE-CD DNSMOS SIG 4.02 #1 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND ROSE-CD ESTOI 0.87 #1 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND ROSE-CD PESQ 3.49 #1 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND ROSE-CD SI-SDR 17.80 #1 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND rose_cd(PESQ ) DNSMOS 3.01 #2 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND rose_cd(PESQ ) DNSMOS BAK 4.29 #2 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND rose_cd(PESQ ) DNSMOS OVRL 3.28 #2 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND rose_cd(PESQ ) DNSMOS SIG 3.52 #2 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND rose_cd(PESQ ) ESTOI 0.83 #2 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND rose_cd(PESQ ) PESQ 3.99 #2 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND rose_cd(PESQ ) PESQ (wb) 3.99 #2 of 2 Archive leaderboard report
Speech Enhancement VoiceBank+DEMAND rose_cd(PESQ ) SI-SDR 0.40 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Consistency ModelsSPEED

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections