{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-comparison-of-adaptation-techniques-and","title":"A Comparison of Adaptation Techniques and Recurrent Neural Network Architectures","arxiv_id":"1807.06441","date":"2018-07-12","proceeding":null,"authors":["Jan Vanek","Josef Michalek","Jan Zelinka","Josef Psutka"],"abstract":"Recently, recurrent neural networks have become state-of-the-art in acoustic\nmodeling for automatic speech recognition. The long short-term memory (LSTM)\nunits are the most popular ones. However, alternative units like gated\nrecurrent unit (GRU) and its modifications outperformed LSTM in some\npublications. In this paper, we compared five neural network (NN) architectures\nwith various adaptation and feature normalization techniques. We have evaluated\nfeature-space maximum likelihood linear regression, five variants of i-vector\nadaptation and two variants of cepstral mean normalization. The most adaptation\nand normalization techniques were developed for feed-forward NNs and, according\nto results in this paper, not all of them worked also with RNNs. For\nexperiments, we have chosen a well known and available TIMIT phone recognition\ntask. The phone recognition is much more sensitive to the quality of AM than\nlarge vocabulary task with a complex language model. Also, we published the\nopen-source scripts to easily replicate the results and to help continue the\ndevelopment.","url_abs":"http://arxiv.org/abs/1807.06441v1","url_pdf":"http://arxiv.org/pdf/1807.06441v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-comparison-of-adaptation-techniques-and","repo_url":"https://github.com/OrcusCZ/NNAcousticModeling","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[{"method_slug":"am","method_name":"AM"},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}