{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/weighted-speech-distortion-losses-for-neural-1","title":"Weighted Speech Distortion Losses for Neural-network-based Real-time Speech Enhancement","arxiv_id":"2001.10601","date":"2020-02-12","proceeding":"IEEE International Conference on Acoustics,  Speech and Signal Processing (ICASSP) 2020 2","authors":[],"abstract":"This paper investigates several aspects of training a RNN (recurrent neural\nnetwork) that impact the objective and subjective quality of enhanced speech\nfor real-time single-channel speech enhancement. Specifically, we focus on a\nRNN that enhances short-time speech spectra on a single-frame-in,\nsingle-frame-out basis, a framework adopted by most classical signal processing\nmethods. We propose two novel mean-squared-error-based learning objectives that\nenable separate control over the importance of speech distortion versus noise\nreduction. The proposed loss functions are evaluated by widely accepted\nobjective quality and intelligibility measures and compared to other\ncompetitive online methods. In addition, we study the impact of feature\nnormalization and varying batch sequence lengths on the objective quality of\nenhanced speech. Finally, we show subjective ratings for the proposed approach\nand a state-of-the-art real-time RNN-based method.","url_abs":"http://arxiv.org/abs/2001.10601v2","url_pdf":"http://arxiv.org/pdf/2001.10601v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"weighted-speech-distortion-losses-for-neural-1","repo_url":"https://github.com/GuillaumeVW/NSNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"weighted-speech-distortion-losses-for-neural-1","repo_url":"https://github.com/MindSpore-scientific/code-6/tree/main/Neural-Network-based-Speech-Enhancement","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"weighted-speech-distortion-losses-for-neural-1","repo_url":"https://github.com/microsoft/DNS-Challenge","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-enhancement-on-deep-noise-suppression","task":"Speech Enhancement","dataset":"Deep Noise Suppression (DNS) Challenge","model":"Proposed (0.35)","rank_in_archive_order":27,"of":36,"metrics":{"PESQ-NB":"2.65","PESQ-WB":"2.65"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}