{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/contaminated-speech-training-methods-for","title":"Contaminated speech training methods for robust DNN-HMM distant speech recognition","arxiv_id":"1710.03538","date":"2017-10-10","proceeding":null,"authors":["Mirco Ravanelli","Maurizio Omologo"],"abstract":"Despite the significant progress made in the last years, state-of-the-art\nspeech recognition technologies provide a satisfactory performance only in the\nclose-talking condition. Robustness of distant speech recognition in adverse\nacoustic conditions, on the other hand, remains a crucial open issue for future\napplications of human-machine interaction. To this end, several advances in\nspeech enhancement, acoustic scene analysis as well as acoustic modeling, have\nrecently contributed to improve the state-of-the-art in the field. One of the\nmost effective approaches to derive a robust acoustic modeling is based on\nusing contaminated speech, which proved helpful in reducing the acoustic\nmismatch between training and testing conditions.\n  In this paper, we revise this classical approach in the context of modern\nDNN-HMM systems, and propose the adoption of three methods, namely, asymmetric\ncontext windowing, close-talk based supervision, and close-talk based\npre-training. The experimental results, obtained using both real and simulated\ndata, show a significant advantage in using these three methods, overall\nproviding a 15% error rate reduction compared to the baseline systems. The same\ntrend in performance is confirmed either using a high-quality training set of\nsmall size, and a large one.","url_abs":"http://arxiv.org/abs/1710.03538v1","url_pdf":"http://arxiv.org/pdf/1710.03538v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"contaminated-speech-training-methods-for","repo_url":"https://github.com/mravanelli/pySpeechRev","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"distant-speech-recognition","task_name":"Distant Speech Recognition"},{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1710.03538","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}