Papers › Conditional Diffusion Probabilistic Model for Speech Enhancement

Conditional Diffusion Probabilistic Model for Speech Enhancement

10 Feb 2022arXiv:2202.05256archive 2025-07-28

Yen-Ju Lu, Zhong-Qiu Wang, Shinji Watanabe, Alexander Richard, Cheng Yu, Yu Tsao

Speech enhancement is a critical component of many user-oriented audio applications, yet current systems still suffer from distorted and unnatural outputs. While generative models have shown strong potential in speech synthesis, they are still lagging behind in speech enhancement. This work leverages recent advances in diffusion probabilistic models, and proposes a novel speech enhancement algorithm that incorporates characteristics of the observed noisy speech signal into the diffusion and reverse processes. More specifically, we propose a generalized formulation of the diffusion probabilistic model named conditional diffusion probabilistic model that, in its reverse process, can adapt to non-Gaussian real noises in the estimated speech signal. In our experiments, we demonstrate strong performance of the proposed approach compared to representative generative models, and investigate the generalization capability of our models to other datasets with noise characteristics unseen during training.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

neillu23/cdiffuse officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Speech EnhancementSpeech Synthesismodel

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Enhancement EARS-WHAM CDiffuSE DNSMOS 2.87 #6 of 6 Archive leaderboard report
Speech Enhancement EARS-WHAM CDiffuSE ESTOI 0.53 #6 of 6 Archive leaderboard report
Speech Enhancement EARS-WHAM CDiffuSE PESQ-WB 1.60 #6 of 6 Archive leaderboard report
Speech Enhancement EARS-WHAM CDiffuSE POLQA 1.81 #6 of 6 Archive leaderboard report
Speech Enhancement EARS-WHAM CDiffuSE SI-SDR 8.35 #6 of 6 Archive leaderboard report
Speech Enhancement EARS-WHAM CDiffuSE SIGMOS 2.08 #6 of 6 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Diffusion

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections