{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/very-deep-convolutional-networks-for-end-to","title":"Very Deep Convolutional Networks for End-to-End Speech Recognition","arxiv_id":"1610.03022","date":"2016-10-10","proceeding":null,"authors":["Yu Zhang","William Chan","Navdeep Jaitly"],"abstract":"Sequence-to-sequence models have shown success in end-to-end speech\nrecognition. However these models have only used shallow acoustic encoder\nnetworks. In our work, we successively train very deep convolutional networks\nto add more expressive power and better generalization for end-to-end ASR\nmodels. We apply network-in-network principles, batch normalization, residual\nconnections and convolutional LSTMs to build very deep recurrent and\nconvolutional structures. Our models exploit the spectral structure in the\nfeature space and add computational depth without overfitting issues. We\nexperiment with the WSJ ASR task and achieve 10.5\\% word error rate without any\ndictionary or language using a 15 layer deep network.","url_abs":"http://arxiv.org/abs/1610.03022v1","url_pdf":"http://arxiv.org/pdf/1610.03022v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"very-deep-convolutional-networks-for-end-to","repo_url":"https://github.com/aaaceo890/Attention","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"very-deep-convolutional-networks-for-end-to","repo_url":"https://github.com/colaprograms/speechify","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1610.03022","atlas_url":"https://app.syntology.ai/?focus=1610.03022","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}