{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/stochastic-training-of-neural-networks-via","title":"Stochastic Training of Neural Networks via Successive Convex Approximations","arxiv_id":"1706.04769","date":"2017-06-15","proceeding":null,"authors":["Simone Scardapane","Paolo Di Lorenzo"],"abstract":"This paper proposes a new family of algorithms for training neural networks\n(NNs). These are based on recent developments in the field of non-convex\noptimization, going under the general name of successive convex approximation\n(SCA) techniques. The basic idea is to iteratively replace the original\n(non-convex, highly dimensional) learning problem with a sequence of (strongly\nconvex) approximations, which are both accurate and simple to optimize.\nDifferently from similar ideas (e.g., quasi-Newton algorithms), the\napproximations can be constructed using only first-order information of the\nneural network function, in a stochastic fashion, while exploiting the overall\nstructure of the learning problem for a faster convergence. We discuss several\nuse cases, based on different choices for the loss function (e.g., squared loss\nand cross-entropy loss), and for the regularization of the NN's weights. We\nexperiment on several medium-sized benchmark problems, and on a large-scale\ndataset involving simulated physical data. The results show how the algorithm\noutperforms state-of-the-art techniques, providing faster convergence to a\nbetter minimum. Additionally, we show how the algorithm can be easily\nparallelized over multiple computational units without hindering its\nperformance. In particular, each computational unit can optimize a tailored\nsurrogate function defined on a randomly assigned subset of the input\nvariables, whose dimension can be selected depending entirely on the available\ncomputational power.","url_abs":"http://arxiv.org/abs/1706.04769v1","url_pdf":"http://arxiv.org/pdf/1706.04769v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"stochastic-training-of-neural-networks-via","repo_url":"https://bitbucket.org/ispamm/sca-optimization-for-neural-networks","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}