{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/is-deeper-better-only-when-shallow-is-good","title":"Is Deeper Better only when Shallow is Good?","arxiv_id":"1903.03488","date":"2019-03-08","proceeding":"NeurIPS 2019 12","authors":["Eran Malach","Shai Shalev-Shwartz"],"abstract":"Understanding the power of depth in feed-forward neural networks is an\nongoing challenge in the field of deep learning theory. While current works\naccount for the importance of depth for the expressive power of\nneural-networks, it remains an open question whether these benefits are\nexploited during a gradient-based optimization process. In this work we explore\nthe relation between expressivity properties of deep networks and the ability\nto train them efficiently using gradient-based algorithms. We give a depth\nseparation argument for distributions with fractal structure, showing that they\ncan be expressed efficiently by deep networks, but not with shallow ones. These\ndistributions have a natural coarse-to-fine structure, and we show that the\nbalance between the coarse and fine details has a crucial effect on whether the\noptimization process is likely to succeed. We prove that when the distribution\nis concentrated on the fine details, gradient-based algorithms are likely to\nfail. Using this result we prove that, at least in some distributions, the\nsuccess of learning deep networks depends on whether the distribution can be\nwell approximated by shallower networks, and we conjecture that this property\nholds in general.","url_abs":"http://arxiv.org/abs/1903.03488v1","url_pdf":"http://arxiv.org/pdf/1903.03488v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"is-deeper-better-only-when-shallow-is-good","repo_url":"https://github.com/emalach/IsDeeperBetter","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"learning-theory","task_name":"Learning Theory"},{"task_slug":"open-question","task_name":"Open-Ended Question Answering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1903.03488","atlas_url":"https://app.syntology.ai/?focus=1903.03488","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}