{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-framework-for-parallel-and-distributed","title":"A Framework for Parallel and Distributed Training of Neural Networks","arxiv_id":"1610.07448","date":"2016-10-24","proceeding":null,"authors":["Simone Scardapane","Paolo Di Lorenzo"],"abstract":"The aim of this paper is to develop a general framework for training neural\nnetworks (NNs) in a distributed environment, where training data is partitioned\nover a set of agents that communicate with each other through a sparse,\npossibly time-varying, connectivity pattern. In such distributed scenario, the\ntraining problem can be formulated as the (regularized) optimization of a\nnon-convex social cost function, given by the sum of local (non-convex) costs,\nwhere each agent contributes with a single error term defined with respect to\nits local dataset. To devise a flexible and efficient solution, we customize a\nrecently proposed framework for non-convex optimization over networks, which\nhinges on a (primal) convexification-decomposition technique to handle\nnon-convexity, and a dynamic consensus procedure to diffuse information among\nthe agents. Several typical choices for the training criterion (e.g., squared\nloss, cross entropy, etc.) and regularization (e.g., $\\ell_2$ norm, sparsity\ninducing penalties, etc.) are included in the framework and explored along the\npaper. Convergence to a stationary solution of the social non-convex problem is\nguaranteed under mild assumptions. Additionally, we show a principled way\nallowing each agent to exploit a possible multi-core architecture (e.g., a\nlocal cloud) in order to parallelize its local optimization step, resulting in\nstrategies that are both distributed (across the agents) and parallel (inside\neach agent) in nature. A comprehensive set of experimental results validate the\nproposed approach.","url_abs":"http://arxiv.org/abs/1610.07448v3","url_pdf":"http://arxiv.org/pdf/1610.07448v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-framework-for-parallel-and-distributed","repo_url":"https://bitbucket.org/ispamm/parallel-and-distributed-neural-networks","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}