Papers › Monaural Speech Enhancement with Complex Convolutional Block Attention Module and...

Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses

3 Feb 2021arXiv:2102.01993archive 2025-07-28

Shengkui Zhao, Trung Hieu Nguyen, Bin Ma

Deep complex U-Net structure and convolutional recurrent network (CRN) structure achieve state-of-the-art performance for monaural speech enhancement. Both deep complex U-Net and CRN are encoder and decoder structures with skip connections, which heavily rely on the representation power of the complex-valued convolutional layers. In this paper, we propose a complex convolutional block attention module (CCBAM) to boost the representation power of the complex-valued convolutional layers by constructing more informative features. The CCBAM is a lightweight and general module which can be easily integrated into any complex-valued convolutional layers. We integrate CCBAM with the deep complex U-Net and CRN to enhance their performance for speech enhancement. We further propose a mixed loss function to jointly optimize the complex models in both time-frequency (TF) domain and time domain. By integrating CCBAM and the mixed loss, we form a new end-to-end (E2E) complex speech enhancement framework. Ablation experiments and objective evaluations show the superior performance of the proposed approaches (https://github.com/modelscope/ClearerVoice-Studio).

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

modelscope/ClearerVoice-Studio officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderSpeech DenoisingSpeech Enhancement

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Enhancement DNS Challenge DCCRN-MC PESQ-NB 3.21 #2 of 5 Archive leaderboard report
Speech Enhancement DNS Challenge DCCRN-M PESQ-NB 3.15 #3 of 5 Archive leaderboard report
Speech Enhancement DNS Challenge DCCRN PESQ-NB 3.04 #4 of 5 Archive leaderboard report
Speech Enhancement Deep Noise Suppression (DNS) Challenge FRCRN PESQ-WB 3.23 #13 of 36 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND D2Former PESQ (wb) 3.43 #14 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND D2Former Para. (M) 0.86 #14 of 42 Archive leaderboard report
Speech Enhancement WSJ0 + DEMAND + RNNoise DCUNet-MC PESQ-NB 3.44 #1 of 3 Archive leaderboard report
Speech Enhancement WSJ0 + DEMAND + RNNoise DCCRN-M PESQ-NB 3.28 #2 of 3 Archive leaderboard report
Speech Enhancement WSJ0 + DEMAND + RNNoise DCUNet PESQ-NB 3.25 #3 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Average PoolingCRNConcatenated Skip ConnectionConvolutionMax PoolingReLUU-Net

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections