{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-learning-without-poor-local-minima","title":"Deep Learning without Poor Local Minima","arxiv_id":"1605.07110","date":"2016-05-23","proceeding":"NeurIPS 2016 12","authors":["Kenji Kawaguchi"],"abstract":"In this paper, we prove a conjecture published in 1989 and also partially\naddress an open problem announced at the Conference on Learning Theory (COLT)\n2015. With no unrealistic assumption, we first prove the following statements\nfor the squared loss function of deep linear neural networks with any depth and\nany widths: 1) the function is non-convex and non-concave, 2) every local\nminimum is a global minimum, 3) every critical point that is not a global\nminimum is a saddle point, and 4) there exist \"bad\" saddle points (where the\nHessian has no negative eigenvalue) for the deeper networks (with more than\nthree layers), whereas there is no bad saddle point for the shallow networks\n(with three layers). Moreover, for deep nonlinear neural networks, we prove the\nsame four statements via a reduction to a deep linear model under the\nindependence assumption adopted from recent work. As a result, we present an\ninstance, for which we can answer the following question: how difficult is it\nto directly train a deep model in theory? It is more difficult than the\nclassical machine learning models (because of the non-convexity), but not too\ndifficult (because of the nonexistence of poor local minima). Furthermore, the\nmathematically proven existence of bad saddle points for deeper models would\nsuggest a possible open problem. We note that even though we have advanced the\ntheoretical foundations of deep learning and non-convex optimization, there is\nstill a gap between theory and practice.","url_abs":"http://arxiv.org/abs/1605.07110v3","url_pdf":"http://arxiv.org/pdf/1605.07110v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-learning-without-poor-local-minima","repo_url":"https://github.com/yijiazh/DFER_Summer2019","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"learning-theory","task_name":"Learning Theory"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1605.07110","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}