Papers › CMGAN: Conformer-Based Metric-GAN for Monaural Speech Enhancement

CMGAN: Conformer-Based Metric-GAN for Monaural Speech Enhancement

22 Sep 2022arXiv:2209.11112archive 2025-07-28

Sherif Abdulatif, Ruizhe Cao, Bin Yang

In this work, we further develop the conformer-based metric generative adversarial network (CMGAN) model for speech enhancement (SE) in the time-frequency (TF) domain. This paper builds on our previous work but takes a more in-depth look by conducting extensive ablation studies on model inputs and architectural design choices. We rigorously tested the generalization ability of the model to unseen noise types and distortions. We have fortified our claims through DNS-MOS measurements and listening tests. Rather than focusing exclusively on the speech denoising task, we extend this work to address the dereverberation and super-resolution tasks. This necessitated exploring various architectural changes, specifically metric discriminator scores and masking techniques. It is essential to highlight that this is among the earliest works that attempted complex TF-domain super-resolution. Our findings show that CMGAN outperforms existing state-of-the-art methods in the three major speech enhancement tasks: denoising, dereverberation, and super-resolution. For example, in the denoising task using the Voice Bank+DEMAND dataset, CMGAN notably exceeded the performance of prior models, attaining a PESQ score of 3.41 and an SSNR of 11.10 dB. Audio samples and CMGAN implementations are available online.

PaperPDFCode

Code

ruizhecao96/cmgan officialmentioned in papermentioned on GitHubpytorchMIT report
SherifAbdulatif/CMGAN officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio Super-ResolutionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)DecoderDenoisingSpeech DenoisingSpeech EnhancementSpeech RecognitionSpeech SeparationSuper-Resolutionspeech-recognition

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio Super-Resolution VCTK Multi-Speaker CMGAN Log-Spectral Distance 0.76 #1 of 7 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND CMGAN CBAK 3.94 #15 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND CMGAN COVL 4.12 #15 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND CMGAN CSIG 4.63 #15 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND CMGAN PESQ (wb) 3.41 #15 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND CMGAN SSNR 11.1 #15 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND CMGAN STOI 96 #15 of 42 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections