{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/d2net-a-denoising-and-dereverberation-network","title":"D²Net: A Denoising and Dereverberation Network Based on Two-branch Encoder and Dual-path Transformer","arxiv_id":null,"date":"2022-11-21","proceeding":"APSIPA ASC 2022 11","authors":["Liusong Wang","Wenbing Wei","Yadong Chen","and Ying Hu"],"abstract":"The simultaneous denoising and dereverberation for single-channel mixture speech under the complicated acoustic environment is considered to be a challengeable task. In this paper, we propose a denoising and dereverberation network named as D²Net in which a two-branch encoder (TBE) is designed to extract and selectively fuse features with different granularity. In addition, we design a global-local dual-path transformer (GLDPT) which introduces the local dense synthesizer attention (LDSA) in the dual-path transformer to improve the perception of local information. We evaluated our proposed D²Net and conducted ablation studies on the VoiceBank+DEMAND and WHAMR! datasets. Meanwhile, we chose three types of data in the WHAMR! dataset to verify the ability of the D²Net on the tasks of denoising-only, dereverberation-only, and simultaneous denoising and dereverberation, respectively. Experimental results show that our proposed model outperforms the comparative models, and all achieve better performance on the tasks of simultaneous denoising and dereverberation, dereverberation-only, and denoising-only, while keeping a small number of network parameters.","url_abs":"https://ieeexplore.ieee.org/abstract/document/9979863","url_pdf":"http://www.apsipa.org/proceedings/2022/APSIPA%202022/ThPM1-2/1570833515.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-enhancement-on-demand","task":"Speech Enhancement","dataset":"VoiceBank + DEMAND","model":"D²Net","rank_in_archive_order":17,"of":42,"metrics":{"CBAK":"3.18","COVL":"3.92","CSIG":"4.63","PESQ (wb)":"3.27","STOI":"96"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}