{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/wpd-an-improved-neural-beamformer-for","title":"WPD++: An Improved Neural Beamformer for Simultaneous Speech Separation and Dereverberation","arxiv_id":"2011.09162","date":"2020-11-18","proceeding":null,"authors":[],"abstract":"This paper aims at eliminating the interfering speakers' speech, additive\nnoise, and reverberation from the noisy multi-talker speech mixture that\nbenefits automatic speech recognition (ASR) backend. While the recently\nproposed Weighted Power minimization Distortionless response (WPD) beamformer\ncan perform separation and dereverberation simultaneously, the noise\ncancellation component still has the potential to progress. We propose an\nimproved neural WPD beamformer called \"WPD++\" by an enhanced beamforming module\nin the conventional WPD and a multi-objective loss function for the joint\ntraining. The beamforming module is improved by utilizing the spatio-temporal\ncorrelation. A multi-objective loss, including the complex spectra domain\nscale-invariant signal-to-noise ratio (C-Si-SNR) and the magnitude domain mean\nsquare error (Mag-MSE), is properly designed to make multiple constraints on\nthe enhanced speech and the desired power of the dry clean signal. Joint\ntraining is conducted to optimize the complex-valued mask estimator and the\nWPD++ beamformer in an end-to-end way. The results show that the proposed WPD++\noutperforms several state-of-the-art beamformers on the enhanced speech quality\nand word error rate (WER) of ASR.","url_abs":"http://arxiv.org/abs/2011.09162v1","url_pdf":"http://arxiv.org/pdf/2011.09162v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"wpd-an-improved-neural-beamformer-for","repo_url":"https://github.com/nateanl/wpd-plus-plus","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-separation","task_name":"Speech Separation"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}