{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-loss-surfaces-of-multilayer-networks","title":"The Loss Surfaces of Multilayer Networks","arxiv_id":"1412.0233","date":"2014-11-30","proceeding":null,"authors":["Anna Choromanska","Mikael Henaff","Michael Mathieu","Gérard Ben Arous","Yann Lecun"],"abstract":"We study the connection between the highly non-convex loss function of a\nsimple model of the fully-connected feed-forward neural network and the\nHamiltonian of the spherical spin-glass model under the assumptions of: i)\nvariable independence, ii) redundancy in network parametrization, and iii)\nuniformity. These assumptions enable us to explain the complexity of the fully\ndecoupled neural network through the prism of the results from random matrix\ntheory. We show that for large-size decoupled networks the lowest critical\nvalues of the random loss function form a layered structure and they are\nlocated in a well-defined band lower-bounded by the global minimum. The number\nof local minima outside that band diminishes exponentially with the size of the\nnetwork. We empirically verify that the mathematical model exhibits similar\nbehavior as the computer simulations, despite the presence of high dependencies\nin real networks. We conjecture that both simulated annealing and SGD converge\nto the band of low critical points, and that all critical points found there\nare local minima of high quality measured by the test error. This emphasizes a\nmajor difference between large- and small-size networks where for the latter\npoor quality local minima have non-zero probability of being recovered.\nFinally, we prove that recovering the global minimum becomes harder as the\nnetwork size increases and that it is in practice irrelevant as global minimum\noften leads to overfitting.","url_abs":"http://arxiv.org/abs/1412.0233v3","url_pdf":"http://arxiv.org/pdf/1412.0233v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-loss-surfaces-of-multilayer-networks","repo_url":"https://github.com/jchunn/Ambition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[],"methods":[{"method_slug":"sgd","method_name":"SGD"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1412.0233","atlas_url":"https://app.syntology.ai/?focus=1412.0233","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}