Papers › CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR

CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR

27 Feb 2025arXiv:2502.20040archive 2025-07-28

Nian Shao, Rui Zhou, Pengyu Wang, Xian Li, Ying Fang, Yujie Yang, Xiaofei Li

In this work, we propose CleanMel, a single-channel Mel-spectrogram denoising and dereverberation network for improving both speech quality and automatic speech recognition (ASR) performance. The proposed network takes as input the noisy and reverberant microphone recording and predicts the corresponding clean Mel-spectrogram. The enhanced Mel-spectrogram can be either transformed to speech waveform with a neural vocoder or directly used for ASR. The proposed network is composed of interleaved cross-band and narrow-band processing in the Mel-frequency domain, for learning the full-band spectral pattern and the narrow-band properties of signals, respectively. Compared to linear-frequency domain or time-domain speech enhancement, the key advantage of Mel-spectrogram enhancement is that Mel-frequency presents speech in a more compact way and thus is easier to learn, which will benefit both speech quality and ASR. Experimental results on four English and one Chinese datasets demonstrate a significant improvement in both speech quality and ASR performance achieved by the proposed model. Code and audio examples of our model are available online in https://audio.westlake.edu.cn/Research/CleanMel.html.

PaperPDFCode

Code

Audio-WestlakeU/CleanMel officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DenoisingSpeech EnhancementSpeech Recognitionspeech-recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Automatic Speech Recognition (ASR) RealMAN CleanMel-L-mask CER 14.4 #1 of 2 Archive leaderboard report
Speech Enhancement RealMAN CleanMel-L-map DNSMOS 3.82 #1 of 2 Archive leaderboard report
Speech Enhancement RealMAN CleanMel-L-map DNSMOS BAK 4.03 #1 of 2 Archive leaderboard report
Speech Enhancement RealMAN CleanMel-L-map DNSMOS OVRL 3.25 #1 of 2 Archive leaderboard report
Speech Enhancement RealMAN CleanMel-L-map DNSMOS SIG 3.55 #1 of 2 Archive leaderboard report
Speech Enhancement RealMAN CleanMel-L-map PESQ-WB 2.10 #1 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections