{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/gradient-energy-matching-for-distributed","title":"Gradient Energy Matching for Distributed Asynchronous Gradient Descent","arxiv_id":"1805.08469","date":"2018-05-22","proceeding":null,"authors":["Joeri Hermans","Gilles Louppe"],"abstract":"Distributed asynchronous SGD has become widely used for deep learning in\nlarge-scale systems, but remains notorious for its instability when increasing\nthe number of workers. In this work, we study the dynamics of distributed\nasynchronous SGD under the lens of Lagrangian mechanics. Using this\ndescription, we introduce the concept of energy to describe the optimization\nprocess and derive a sufficient condition ensuring its stability as long as the\ncollective energy induced by the active workers remains below the energy of a\ntarget synchronous process. Making use of this criterion, we derive a stable\ndistributed asynchronous optimization procedure, GEM, that estimates and\nmaintains the energy of the asynchronous system below or equal to the energy of\nsequential SGD with momentum. Experimental results highlight the stability and\nspeedup of GEM compared to existing schemes, even when scaling to one hundred\nasynchronous workers. Results also indicate better generalization compared to\nthe targeted SGD with momentum.","url_abs":"http://arxiv.org/abs/1805.08469v1","url_pdf":"http://arxiv.org/pdf/1805.08469v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"gradient-energy-matching-for-distributed","repo_url":"https://github.com/montefiore-ai/gradient-energy-matching","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"gradient-energy-matching-for-distributed","repo_url":"https://github.com/vlimant/NNLO","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[],"methods":[{"method_slug":"sgd","method_name":"SGD"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}