Papers › Perceptual Loss based Speech Denoising with an ensemble of Audio Pattern Recognition...

Perceptual Loss based Speech Denoising with an ensemble of Audio Pattern Recognition and Self-Supervised Models

22 Oct 2020arXiv:2010.11860archive 2025-07-28

Deep learning based speech denoising still suffers from the challenge of improving perceptual quality of enhanced signals. We introduce a generalized framework called Perceptual Ensemble Regularization Loss (PERL) built on the idea of perceptual losses. Perceptual loss discourages distortion to certain speech properties and we analyze it using six large-scale pre-trained models: speaker classification, acoustic model, speaker embedding, emotion classification, and two self-supervised speech encoders (PASE+, wav2vec 2.0). We first build a strong baseline (w/o PERL) using Conformer Transformer Networks on the popular enhancement benchmark called VCTK-DEMAND. Using auxiliary models one at a time, we find acoustic event and self-supervised model PASE+ to be most effective. Our best model (PERL-AE) only uses acoustic event model (utilizing AudioSet) to outperform state-of-the-art methods on major perceptual metrics. To explore if denoising can leverage full framework, we use all networks but find that our seven-loss formulation suffers from the challenges of Multi-Task Learning. Finally, we report a critical observation that state-of-the-art Multi-Task weight learning methods cannot outperform hand tuning, perhaps due to challenges of domain mismatch and weak complementarity of losses.

PaperPDFConference PDFCode

Code

saurabh-kataria/PERL-samples mentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DenoisingEmotion ClassificationMulti-Task LearningSpeech Denoising

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Enhancement VoiceBank + DEMAND PERL-AE CBAK 3.53 #24 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND PERL-AE COVL 3.83 #24 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND PERL-AE CSIG 4.43 #24 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND PERL-AE PESQ (wb) 3.17 #24 of 42 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

PASE+

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections