{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/do-deep-nets-really-need-to-be-deep","title":"Do Deep Nets Really Need to be Deep?","arxiv_id":"1312.6184","date":"2013-12-21","proceeding":"NeurIPS 2014 12","authors":["Lei Jimmy Ba","Rich Caruana"],"abstract":"Currently, deep neural networks are the state of the art on problems such as\nspeech recognition and computer vision. In this extended abstract, we show that\nshallow feed-forward networks can learn the complex functions previously\nlearned by deep nets and achieve accuracies previously only achievable with\ndeep models. Moreover, in some cases the shallow neural nets can learn these\ndeep functions using a total number of parameters similar to the original deep\nmodel. We evaluate our method on the TIMIT phoneme recognition task and are\nable to train shallow fully-connected nets that perform similarly to complex,\nwell-engineered, deep convolutional architectures. Our success in training\nshallow neural nets to mimic deeper models suggests that there probably exist\nbetter algorithms for training shallow feed-forward nets than those currently\navailable.","url_abs":"http://arxiv.org/abs/1312.6184v7","url_pdf":"http://arxiv.org/pdf/1312.6184v7.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"do-deep-nets-really-need-to-be-deep","repo_url":"https://github.com/jchen98/compression","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"do-deep-nets-really-need-to-be-deep","repo_url":"https://github.com/peta78/linear-regression-voting-and-statistics","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"gone","observed_at":"2026-09-18","how":"tree_404+repo_404"}}],"tasks":[{"task_slug":"phoneme-recognition","task_name":"Phoneme Recognition"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1312.6184","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}