{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/collapse-of-deep-and-narrow-neural-nets","title":"Collapse of Deep and Narrow Neural Nets","arxiv_id":"1808.04947","date":"2018-08-15","proceeding":"ICLR 2019 5","authors":["Lu Lu","Yanhui Su","George Em. Karniadakis"],"abstract":"Recent theoretical work has demonstrated that deep neural networks have\nsuperior performance over shallow networks, but their training is more\ndifficult, e.g., they suffer from the vanishing gradient problem. This problem\ncan be typically resolved by the rectified linear unit (ReLU) activation.\nHowever, here we show that even for such activation, deep and narrow neural\nnetworks (NNs) will converge to erroneous mean or median states of the target\nfunction depending on the loss with high probability. Deep and narrow NNs are\nencountered in solving partial differential equations with high-order\nderivatives. We demonstrate this collapse of such NNs both numerically and\ntheoretically, and provide estimates of the probability of collapse. We also\nconstruct a diagram of a safe region for designing NNs that avoid the collapse\nto erroneous states. Finally, we examine different ways of initialization and\nnormalization that may avoid the collapse problem. Asymmetric initializations\nmay reduce the probability of collapse but do not totally eliminate it.","url_abs":"http://arxiv.org/abs/1808.04947v2","url_pdf":"http://arxiv.org/pdf/1808.04947v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"collapse-of-deep-and-narrow-neural-nets","repo_url":"https://github.com/ericpts/vae-res","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1808.04947","atlas_url":"https://app.syntology.ai/?focus=1808.04947","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}