{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-from-past-mistakes-improving","title":"Learning from Past Mistakes: Improving Automatic Speech Recognition Output via Noisy-Clean Phrase Context Modeling","arxiv_id":"1802.02607","date":"2018-02-07","proceeding":null,"authors":["Prashanth Gurunath Shivakumar","Haoqi Li","Kevin Knight","Panayiotis Georgiou"],"abstract":"Automatic speech recognition (ASR) systems often make unrecoverable errors\ndue to subsystem pruning (acoustic, language and pronunciation models); for\nexample pruning words due to acoustics using short-term context, prior to\nrescoring with long-term context based on linguistics. In this work we model\nASR as a phrase-based noisy transformation channel and propose an error\ncorrection system that can learn from the aggregate errors of all the\nindependent modules constituting the ASR and attempt to invert those. The\nproposed system can exploit long-term context using a neural network language\nmodel and can better choose between existing ASR output possibilities as well\nas re-introduce previously pruned or unseen (out-of-vocabulary) phrases. It\nprovides corrections under poorly performing ASR conditions without degrading\nany accurate transcriptions; such corrections are greater on top of\nout-of-domain and mismatched data ASR. Our system consistently provides\nimprovements over the baseline ASR, even when baseline is further optimized\nthrough recurrent neural network language model rescoring. This demonstrates\nthat any ASR improvements can be exploited independently and that our proposed\nsystem can potentially still provide benefits on highly optimized ASR. Finally,\nwe present an extensive analysis of the type of errors corrected by our system.","url_abs":"http://arxiv.org/abs/1802.02607v2","url_pdf":"http://arxiv.org/pdf/1802.02607v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-from-past-mistakes-improving","repo_url":"https://github.com/cassandra-lehmann/ensemble_methods_ASR_transcripts","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[{"method_slug":"pruning","method_name":"Pruning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1802.02607","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}