Papers › Speech Enhancement and Dereverberation with Diffusion-based Generative Models

Speech Enhancement and Dereverberation with Diffusion-based Generative Models

11 Aug 2022IEEE/ACM Transactions on Audio, Speech, and Language Processing 2023 6arXiv:2208.05830archive 2025-07-28

Julius Richter, Simon Welker, Jean-Marie Lemercier, Bunlong Lay, Timo Gerkmann

In this work, we build upon our previous publication and use diffusion-based generative models for speech enhancement. We present a detailed overview of the diffusion process that is based on a stochastic differential equation and delve into an extensive theoretical examination of its implications. Opposed to usual conditional generation tasks, we do not start the reverse process from pure Gaussian noise but from a mixture of noisy speech and Gaussian noise. This matches our forward process which moves from clean speech to noisy speech by including a drift term. We show that this procedure enables using only 30 diffusion steps to generate high-quality clean speech estimates. By adapting the network architecture, we are able to significantly improve the speech enhancement performance, indicating that the network, rather than the formalism, was the main limitation of our original approach. In an extensive cross-dataset evaluation, we show that the improved method can compete with recent discriminative models and achieves better generalization when evaluating on a different corpus than used for training. We complement the results with an instrumental evaluation using real-world noisy recordings and a listening experiment, in which our proposed method is rated best. Examining different sampler configurations for solving the reverse process allows us to balance the performance and computational speed of the proposed method. Moreover, we show that the proposed method is also suitable for dereverberation and thus not limited to additive background noise removal. Code and audio examples are available online, see https://github.com/sp-uhh/sgmse

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

sp-uhh/sgmse officialmentioned in papermentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Speech DereverberationSpeech Enhancement

Datasets

Introduced by this paper, per the archive.

Reverb-WSJ0

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Dereverberation EARS-Reverb SGMSE+ ESTOI 0.85 #1 of 1 Archive leaderboard report
Speech Dereverberation EARS-Reverb SGMSE+ MOS Reverb 4.73 #1 of 1 Archive leaderboard report
Speech Dereverberation EARS-Reverb SGMSE+ PESQ-WB 3.03 #1 of 1 Archive leaderboard report
Speech Dereverberation EARS-Reverb SGMSE+ SI-SDR 5.79 #1 of 1 Archive leaderboard report
Speech Dereverberation EARS-Reverb SGMSE+ SIGMOS 3.49 #1 of 1 Archive leaderboard report
Speech Enhancement EARS-WHAM SGMSE+ DNSMOS 3.88 #2 of 6 Archive leaderboard report
Speech Enhancement EARS-WHAM SGMSE+ ESTOI 0.73 #2 of 6 Archive leaderboard report
Speech Enhancement EARS-WHAM SGMSE+ PESQ-WB 2.50 #2 of 6 Archive leaderboard report
Speech Enhancement EARS-WHAM SGMSE+ POLQA 3.40 #2 of 6 Archive leaderboard report
Speech Enhancement EARS-WHAM SGMSE+ SI-SDR 16.78 #2 of 6 Archive leaderboard report
Speech Enhancement EARS-WHAM SGMSE+ SIGMOS 3.41 #2 of 6 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND SGMSE+ (Diffusion Model) PESQ (wb) 2.93 #38 of 42 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

DiffusionSPEEDTest

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections