{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-comparison-of-techniques-for-language-model","title":"A Comparison of Techniques for Language Model Integration in Encoder-Decoder Speech Recognition","arxiv_id":"1807.10857","date":"2018-07-27","proceeding":null,"authors":["Shubham Toshniwal","Anjuli Kannan","Chung-Cheng Chiu","Yonghui Wu","Tara N. Sainath","Karen Livescu"],"abstract":"Attention-based recurrent neural encoder-decoder models present an elegant\nsolution to the automatic speech recognition problem. This approach folds the\nacoustic model, pronunciation model, and language model into a single network\nand requires only a parallel corpus of speech and text for training. However,\nunlike in conventional approaches that combine separate acoustic and language\nmodels, it is not clear how to use additional (unpaired) text. While there has\nbeen previous work on methods addressing this problem, a thorough comparison\namong methods is still lacking. In this paper, we compare a suite of past\nmethods and some of our own proposed methods for using unpaired text data to\nimprove encoder-decoder models. For evaluation, we use the medium-sized\nSwitchboard data set and the large-scale Google voice search and dictation data\nsets. Our results confirm the benefits of using unpaired text across a range of\nmethods and data sets. Surprisingly, for first-pass decoding, the rather simple\napproach of shallow fusion performs best across data sets. However, for Google\ndata sets we find that cold fusion has a lower oracle error rate and\noutperforms other approaches after second-pass rescoring on the Google voice\nsearch data set.","url_abs":"http://arxiv.org/abs/1807.10857v2","url_pdf":"http://arxiv.org/pdf/1807.10857v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-comparison-of-techniques-for-language-model","repo_url":"https://github.com/pwc-1/Paper-10/tree/main/speech_encoder_decoder","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1807.10857","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}