{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rank-1-constrained-multichannel-wiener-filter","title":"Rank-1 Constrained Multichannel Wiener Filter for Speech Recognition in Noisy Environments","arxiv_id":"1707.00201","date":"2017-07-01","proceeding":null,"authors":["Ziteng Wang","Emmanuel Vincent","Romain Serizel","Yonghong Yan"],"abstract":"Multichannel linear filters, such as the Multichannel Wiener Filter (MWF) and\nthe Generalized Eigenvalue (GEV) beamformer are popular signal processing\ntechniques which can improve speech recognition performance. In this paper, we\npresent an experimental study on these linear filters in a specific speech\nrecognition task, namely the CHiME-4 challenge, which features real recordings\nin multiple noisy environments. Specifically, the rank-1 MWF is employed for\nnoise reduction and a new constant residual noise power constraint is derived\nwhich enhances the recognition performance. To fulfill the underlying rank-1\nassumption, the speech covariance matrix is reconstructed based on eigenvectors\nor generalized eigenvectors. Then the rank-1 constrained MWF is evaluated with\nalternative multichannel linear filters under the same framework, which\ninvolves a Bidirectional Long Short-Term Memory (BLSTM) network for mask\nestimation. The proposed filter outperforms alternative ones, leading to a 40%\nrelative Word Error Rate (WER) reduction compared with the baseline Weighted\nDelay and Sum (WDAS) beamformer on the real test set, and a 15% relative WER\nreduction compared with the GEV-BAN method. The results also suggest that the\nspeech recognition accuracy correlates more with the Mel-frequency cepstral\ncoefficients (MFCC) feature variance than with the noise reduction or the\nspeech distortion level.","url_abs":"http://arxiv.org/abs/1707.00201v2","url_pdf":"http://arxiv.org/pdf/1707.00201v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rank-1-constrained-multichannel-wiener-filter","repo_url":"https://github.com/ZitengWang/nn_mask","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}