{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-fsmn-for-large-vocabulary-continuous","title":"Deep-FSMN for Large Vocabulary Continuous Speech Recognition","arxiv_id":"1803.05030","date":"2018-03-04","proceeding":null,"authors":["Shiliang Zhang","Ming Lei","Zhijie Yan","Li-Rong Dai"],"abstract":"In this paper, we present an improved feedforward sequential memory networks\n(FSMN) architecture, namely Deep-FSMN (DFSMN), by introducing skip connections\nbetween memory blocks in adjacent layers. These skip connections enable the\ninformation flow across different layers and thus alleviate the gradient\nvanishing problem when building very deep structure. As a result, DFSMN\nsignificantly benefits from these skip connections and deep structure. We have\ncompared the performance of DFSMN to BLSTM both with and without lower frame\nrate (LFR) on several large speech recognition tasks, including English and\nMandarin. Experimental results shown that DFSMN can consistently outperform\nBLSTM with dramatic gain, especially trained with LFR using CD-Phone as\nmodeling units. In the 2000 hours Fisher (FSH) task, the proposed DFSMN can\nachieve a word error rate of 9.4% by purely using the cross-entropy criterion\nand decoding with a 3-gram language model, which achieves a 1.5% absolute\nimprovement compared to the BLSTM. In a 20000 hours Mandarin recognition task,\nthe LFR trained DFSMN can achieve more than 20% relative improvement compared\nto the LFR trained BLSTM. Moreover, we can easily design the lookahead filter\norder of the memory blocks in DFSMN to control the latency for real-time\napplications.","url_abs":"http://arxiv.org/abs/1803.05030v1","url_pdf":"http://arxiv.org/pdf/1803.05030v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-fsmn-for-large-vocabulary-continuous","repo_url":"https://github.com/yangxueruivs/DFSMN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1803.05030","atlas_url":"https://app.syntology.ai/?focus=1803.05030","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}