{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/very-deep-convolutional-neural-networks-for-1","title":"Very Deep Convolutional Neural Networks for Robust Speech Recognition","arxiv_id":"1610.00277","date":"2016-10-02","proceeding":null,"authors":["Yanmin Qian","Philip C. Woodland"],"abstract":"This paper describes the extension and optimization of our previous work on\nvery deep convolutional neural networks (CNNs) for effective recognition of\nnoisy speech in the Aurora 4 task. The appropriate number of convolutional\nlayers, the sizes of the filters, pooling operations and input feature maps are\nall modified: the filter and pooling sizes are reduced and dimensions of input\nfeature maps are extended to allow adding more convolutional layers.\nFurthermore appropriate input padding and input feature map selection\nstrategies are developed. In addition, an adaptation framework using joint\ntraining of very deep CNN with auxiliary features i-vector and fMLLR features\nis developed. These modifications give substantial word error rate reductions\nover the standard CNN used as baseline. Finally the very deep CNN is combined\nwith an LSTM-RNN acoustic model and it is shown that state-level weighted log\nlikelihood score combination in a joint acoustic model decoding scheme is very\neffective. On the Aurora 4 task, the very deep CNN achieves a WER of 8.81%,\nfurther 7.99% with auxiliary feature joint training, and 7.09% with LSTM-RNN\njoint decoding.","url_abs":"http://arxiv.org/abs/1610.00277v1","url_pdf":"http://arxiv.org/pdf/1610.00277v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"very-deep-convolutional-neural-networks-for-1","repo_url":"https://github.com/Anustup900/Tensorflow-Speech-Recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"very-deep-convolutional-neural-networks-for-1","repo_url":"https://github.com/jayant766/MIDAS-IIITD","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"robust-speech-recognition","task_name":"Robust Speech Recognition"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}