{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/recurrent-dnns-and-its-ensembles-on-the-timit","title":"Recurrent DNNs and its Ensembles on the TIMIT Phone Recognition Task","arxiv_id":"1806.07186","date":"2018-06-19","proceeding":null,"authors":["Jan Vanek","Josef Michalek","Josef Psutka"],"abstract":"In this paper, we have investigated recurrent deep neural networks (DNNs) in\ncombination with regularization techniques as dropout, zoneout, and\nregularization post-layer. As a benchmark, we chose the TIMIT phone recognition\ntask due to its popularity and broad availability in the community. It also\nsimulates a low-resource scenario that is helpful in minor languages. Also, we\nprefer the phone recognition task because it is much more sensitive to an\nacoustic model quality than a large vocabulary continuous speech recognition\ntask. In recent years, recurrent DNNs pushed the error rates in automatic\nspeech recognition down. But, there was no clear winner in proposed\narchitectures. The dropout was used as the regularization technique in most\ncases, but combination with other regularization techniques together with model\nensembles was omitted. However, just an ensemble of recurrent DNNs performed\nbest and achieved an average phone error rate from 10 experiments 14.84 %\n(minimum 14.69 %) on core test set that is slightly lower then the\nbest-published PER to date, according to our knowledge. Finally, in contrast of\nthe most papers, we published the open-source scripts to easily replicate the\nresults and to help continue the development.","url_abs":"http://arxiv.org/abs/1806.07186v1","url_pdf":"http://arxiv.org/pdf/1806.07186v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"recurrent-dnns-and-its-ensembles-on-the-timit","repo_url":"https://github.com/OrcusCZ/NNAcousticModeling","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[{"method_slug":"dropout","method_name":"Dropout"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}