{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/asynchrony-begets-momentum-with-an","title":"Asynchrony begets Momentum, with an Application to Deep Learning","arxiv_id":"1605.09774","date":"2016-05-31","proceeding":null,"authors":["Ioannis Mitliagkas","Ce Zhang","Stefan Hadjis","Christopher Ré"],"abstract":"Asynchronous methods are widely used in deep learning, but have limited\ntheoretical justification when applied to non-convex problems. We show that\nrunning stochastic gradient descent (SGD) in an asynchronous manner can be\nviewed as adding a momentum-like term to the SGD iteration. Our result does not\nassume convexity of the objective function, so it is applicable to deep\nlearning systems. We observe that a standard queuing model of asynchrony\nresults in a form of momentum that is commonly used by deep learning\npractitioners. This forges a link between queuing theory and asynchrony in deep\nlearning systems, which could be useful for systems builders. For convolutional\nneural networks, we experimentally validate that the degree of asynchrony\ndirectly correlates with the momentum, confirming our main result. An important\nimplication is that tuning the momentum parameter is important when considering\ndifferent levels of asynchrony. We assert that properly tuned momentum reduces\nthe number of steps required for convergence. Finally, our theory suggests new\nways of counteracting the adverse effects of asynchrony: a simple mechanism\nlike using negative algorithmic momentum can improve performance under high\nasynchrony. Since asynchronous methods have better hardware efficiency, this\nresult may shed light on when asynchronous execution is more efficient for deep\nlearning systems.","url_abs":"http://arxiv.org/abs/1605.09774v2","url_pdf":"http://arxiv.org/pdf/1605.09774v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"asynchrony-begets-momentum-with-an","repo_url":"https://github.com/JoeriHermans/dist-keras","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"asynchrony-begets-momentum-with-an","repo_url":"https://github.com/cerndb/dist-keras","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"asynchrony-begets-momentum-with-an","repo_url":"https://github.com/valentinchelle/kerasOnSpark","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"}],"methods":[{"method_slug":"sgd","method_name":"SGD"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1605.09774","atlas_url":"https://app.syntology.ai/?focus=1605.09774","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}