{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improving-lstm-ctc-based-asr-performance-in","title":"Improving LSTM-CTC based ASR performance in domains with limited training data","arxiv_id":"1707.00722","date":"2017-07-03","proceeding":null,"authors":["Jayadev Billa"],"abstract":"This paper addresses the observed performance gap between automatic speech\nrecognition (ASR) systems based on Long Short Term Memory (LSTM) neural\nnetworks trained with the connectionist temporal classification (CTC) loss\nfunction and systems based on hybrid Deep Neural Networks (DNNs) trained with\nthe cross entropy (CE) loss function on domains with limited data. We step\nthrough a number of experiments that show incremental improvements on a\nbaseline EESEN toolkit based LSTM-CTC ASR system trained on the Librispeech\n100hr (train-clean-100) corpus. Our results show that with effective\ncombination of data augmentation and regularization, a LSTM-CTC based system\ncan exceed the performance of a strong Kaldi based baseline trained on the same\ndata.","url_abs":"http://arxiv.org/abs/1707.00722v2","url_pdf":"http://arxiv.org/pdf/1707.00722v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improving-lstm-ctc-based-asr-performance-in","repo_url":"https://github.com/jb1999/eesen","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}