Papers › A Modulation-Domain Loss for Neural-Network-based Real-time Speech Enhancement

A Modulation-Domain Loss for Neural-Network-based Real-time Speech Enhancement

15 Feb 2021arXiv:2102.07330archive 2025-07-28

We describe a modulation-domain loss function for deep-learning-based speech enhancement systems. Learnable spectro-temporal receptive fields (STRFs) were adapted to optimize for a speaker identification task. The learned STRFs were then used to calculate a weighted mean-squared error (MSE) in the modulation domain for training a speech enhancement system. Experiments showed that adding the modulation-domain MSE to the MSE in the spectro-temporal domain substantially improved the objective prediction of speech quality and intelligibility for real-time speech enhancement systems without incurring additional computation during inference.

PaperPDFConference PDFCode

Code

tvuong123/ModulationDomainLoss officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Speaker IdentificationSpeech DenoisingSpeech Enhancement

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Enhancement DNS Challenge RNN-Modulation PESQ-WB 2.75 #5 of 5 Archive leaderboard report
Speech Enhancement Deep Noise Suppression (DNS) Challenge RNN-Modulation PESQ-WB 2.75 #24 of 36 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND real-time-GRU PESQ (wb) 2.82 #41 of 42 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections