{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-speech-recognition-using-lattice","title":"End-to-end speech recognition using lattice-free MMI","arxiv_id":null,"date":"2018-09-06","proceeding":"Interspeech 2018 2018 9","authors":["Hossein Hadian","Hossein Sameti","Daniel Povey","Sanjeev Khudanpur"],"abstract":"We present our work on end-to-end training of acoustic models\r\nusing the lattice-free maximum mutual information (LF-MMI)\r\nobjective function in the context of hidden Markov models.\r\nBy end-to-end training, we mean flat-start training of a single\r\nDNN in one stage without using any previously trained models,\r\nforced alignments, or building state-tying decision trees. We\r\nuse full biphones to enable context-dependent modeling without trees, and show that our end-to-end LF-MMI approach can\r\nachieve comparable results to regular LF-MMI on well-known\r\nlarge vocabulary tasks. We also compare with other end-to-end\r\nmethods such as CTC in character-based and lexicon-free settings and show 5 to 25 percent relative reduction in word error rates on different large vocabulary tasks while using significantly smaller models.","url_abs":"https://www.isca-speech.org/archive/Interspeech_2018/abstracts/1423.html","url_pdf":"https://www.danielpovey.com/files/2018_interspeech_end2end.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-recognition-on-switchboard-300hr","task":"Speech Recognition","dataset":"Switchboard (300hr)","model":"End-to-end LF-MMI","rank_in_archive_order":1,"of":1,"metrics":{"Word Error Rate (WER)":"9.3"},"uses_additional_data":false},{"leaderboard":"/sota/speech-recognition-on-wsj-eval92","task":"Speech Recognition","dataset":"WSJ eval92","model":"End-to-end LF-MMI","rank_in_archive_order":7,"of":17,"metrics":{"Word Error Rate (WER)":"3.0"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}