{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exploring-the-best-loss-function-for-dnn-1","title":"Exploring the Best Loss Function for DNN-Based Low-latency Speech Enhancement with Temporal Convolutional Networks","arxiv_id":"2005.11611","date":"2020-08-20","proceeding":"Interspeech 2020 5","authors":[],"abstract":"Recently, deep neural networks (DNNs) have been successfully used for speech\nenhancement, and DNN-based speech enhancement is becoming an attractive\nresearch area. While time-frequency masking based on the short-time Fourier\ntransform (STFT) has been widely used for DNN-based speech enhancement over the\nlast years, time domain methods such as the time-domain audio separation\nnetwork (TasNet) have also been proposed. The most suitable method depends on\nthe scale of the dataset and the type of task. In this paper, we explore the\nbest speech enhancement algorithm on two different datasets. We propose a\nSTFT-based method and a loss function using problem-agnostic speech encoder\n(PASE) features to improve subjective quality for the smaller dataset. Our\nproposed methods are effective on the Voice Bank + DEMAND dataset and compare\nfavorably to other state-of-the-art methods. We also implement a low-latency\nversion of TasNet, which we submitted to the DNS Challenge and made public by\nopen-sourcing it. Our model achieves excellent performance on the DNS Challenge\ndataset.","url_abs":"http://arxiv.org/abs/2005.11611v3","url_pdf":"http://arxiv.org/pdf/2005.11611v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-dereverberation-on-deep-noise","task":"Speech Dereverberation","dataset":"Deep Noise Suppression (DNS) Challenge","model":"Conv-TasNet-SNR","rank_in_archive_order":1,"of":2,"metrics":{"PESQ":"2.75","ΔPESQ":"0.93"},"uses_additional_data":false},{"leaderboard":"/sota/speech-dereverberation-on-deep-noise","task":"Speech Dereverberation","dataset":"Deep Noise Suppression (DNS) Challenge","model":"Noisy/unprocessed","rank_in_archive_order":2,"of":2,"metrics":{"PESQ":"1.82"},"uses_additional_data":false},{"leaderboard":"/sota/speech-enhancement-on-deep-noise-suppression","task":"Speech Enhancement","dataset":"Deep Noise Suppression (DNS) Challenge","model":"Conv-TasNet-SNR","rank_in_archive_order":25,"of":36,"metrics":{"PESQ-WB":"2.73"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2005.11611","atlas_url":"https://app.syntology.ai/?focus=2005.11611","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}