{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/analyzing-the-impact-of-speaker-localization","title":"Analyzing the impact of speaker localization errors on speech separation for automatic speech recognition","arxiv_id":"1910.11114","date":"2019-10-24","proceeding":null,"authors":[],"abstract":"We investigate the effect of speaker localization on the performance of\nspeech recognition systems in a multispeaker, multichannel environment. Given\nthe speaker location information, speech separation is performed in three\nstages. In the first stage, a simple delay-and-sum (DS) beamformer is used to\nenhance the signal impinging from the speaker location which is then used to\nestimate a time-frequency mask corresponding to the localized speaker using a\nneural network. This mask is used to compute the second order statistics and to\nderive an adaptive beamformer in the third stage. We generated a multichannel,\nmultispeaker, reverberated, noisy dataset inspired from the well studied\nWSJ0-2mix and study the performance of the proposed pipeline in terms of the\nword error rate (WER). An average WER of $29.4$% was achieved using the ground\ntruth localization information and $42.4$% using the localization information\nestimated via GCC-PHAT. The signal-to-interference ratio (SIR) between the\nspeakers has a higher impact on the ASR performance, to the extent of reducing\nthe WER by $59$% relative for a SIR increase of $15$ dB. By contrast,\nincreasing the spatial distance to $50^\\circ$ or more improves the WER by $23$%\nrelative only","url_abs":"http://arxiv.org/abs/1910.11114v1","url_pdf":"http://arxiv.org/pdf/1910.11114v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"analyzing-the-impact-of-speaker-localization","repo_url":"https://github.com/sunits/Reverberated_WSJ_2MIX","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-separation","task_name":"Speech Separation"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[{"slug":"kinect-wsj","name":"Kinect-WSJ","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}