{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/atomo-communication-efficient-learning-via","title":"ATOMO: Communication-efficient Learning via Atomic Sparsification","arxiv_id":"1806.04090","date":"2018-06-11","proceeding":"NeurIPS 2018 12","authors":["Hongyi Wang","Scott Sievert","Zachary Charles","Shengchao Liu","Stephen Wright","Dimitris Papailiopoulos"],"abstract":"Distributed model training suffers from communication overheads due to\nfrequent gradient updates transmitted between compute nodes. To mitigate these\noverheads, several studies propose the use of sparsified stochastic gradients.\nWe argue that these are facets of a general sparsification method that can\noperate on any possible atomic decomposition. Notable examples include\nelement-wise, singular value, and Fourier decompositions. We present ATOMO, a\ngeneral framework for atomic sparsification of stochastic gradients. Given a\ngradient, an atomic decomposition, and a sparsity budget, ATOMO gives a random\nunbiased sparsification of the atoms minimizing variance. We show that recent\nmethods such as QSGD and TernGrad are special cases of ATOMO and that\nsparsifiying the singular value decomposition of neural networks gradients,\nrather than their coordinates, can lead to significantly faster distributed\ntraining.","url_abs":"http://arxiv.org/abs/1806.04090v3","url_pdf":"http://arxiv.org/pdf/1806.04090v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"atomo-communication-efficient-learning-via","repo_url":"https://github.com/hwang595/ATOMO","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1806.04090","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}