{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/exponentially-vanishing-sub-optimal-local","title":"Exponentially vanishing sub-optimal local minima in multilayer neural networks","arxiv_id":"1702.05777","date":"2017-02-19","proceeding":"ICLR 2018 1","authors":["Daniel Soudry","Elad Hoffer"],"abstract":"Background: Statistical mechanics results (Dauphin et al. (2014); Choromanska\net al. (2015)) suggest that local minima with high error are exponentially rare\nin high dimensions. However, to prove low error guarantees for Multilayer\nNeural Networks (MNNs), previous works so far required either a heavily\nmodified MNN model or training method, strong assumptions on the labels (e.g.,\n\"near\" linear separability), or an unrealistic hidden layer with\n$\\Omega\\left(N\\right)$ units.\n  Results: We examine a MNN with one hidden layer of piecewise linear units, a\nsingle output, and a quadratic loss. We prove that, with high probability in\nthe limit of $N\\rightarrow\\infty$ datapoints, the volume of differentiable\nregions of the empiric loss containing sub-optimal differentiable local minima\nis exponentially vanishing in comparison with the same volume of global minima,\ngiven standard normal input of dimension\n$d_{0}=\\tilde{\\Omega}\\left(\\sqrt{N}\\right)$, and a more realistic number of\n$d_{1}=\\tilde{\\Omega}\\left(N/d_{0}\\right)$ hidden units. We demonstrate our\nresults numerically: for example, $0\\%$ binary classification training error on\nCIFAR with only $N/d_{0}\\approx 16$ hidden neurons.","url_abs":"http://arxiv.org/abs/1702.05777v5","url_pdf":"http://arxiv.org/pdf/1702.05777v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"exponentially-vanishing-sub-optimal-local","repo_url":"https://github.com/MNNsMinima/Paper","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"gone","observed_at":"2026-09-18","how":"tree_404+repo_404"}}],"tasks":[{"task_slug":"binary-classification","task_name":"Binary Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1702.05777","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}