{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/signsgd-compressed-optimisation-for-non","title":"signSGD: Compressed Optimisation for Non-Convex Problems","arxiv_id":"1802.04434","date":"2018-02-13","proceeding":"ICML 2018 7","authors":["Jeremy Bernstein","Yu-Xiang Wang","Kamyar Azizzadenesheli","Anima Anandkumar"],"abstract":"Training large neural networks requires distributing learning across multiple\nworkers, where the cost of communicating gradients can be a significant\nbottleneck. signSGD alleviates this problem by transmitting just the sign of\neach minibatch stochastic gradient. We prove that it can get the best of both\nworlds: compressed gradients and SGD-level convergence rate. The relative\n$\\ell_1/\\ell_2$ geometry of gradients, noise and curvature informs whether\nsignSGD or SGD is theoretically better suited to a particular problem. On the\npractical side we find that the momentum counterpart of signSGD is able to\nmatch the accuracy and convergence speed of Adam on deep Imagenet models. We\nextend our theory to the distributed setting, where the parameter server uses\nmajority vote to aggregate gradient signs from each worker enabling 1-bit\ncompression of worker-server communication in both directions. Using a theorem\nby Gauss we prove that majority vote can achieve the same reduction in variance\nas full precision distributed SGD. Thus, there is great promise for sign-based\noptimisation schemes to achieve fast communication and fast convergence. Code\nto reproduce experiments is to be found at https://github.com/jxbz/signSGD .","url_abs":"http://arxiv.org/abs/1802.04434v3","url_pdf":"http://arxiv.org/pdf/1802.04434v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"signsgd-compressed-optimisation-for-non","repo_url":"https://github.com/jxbz/signSGD","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"signsgd-compressed-optimisation-for-non","repo_url":"https://github.com/bojone/tiger","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"signsgd-compressed-optimisation-for-non","repo_url":"https://github.com/jasonakoun/signsgd-fault-tolerance","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"signsgd-compressed-optimisation-for-non","repo_url":"https://github.com/scottjiao/Gradient-Compression-Methods","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"signsgd-compressed-optimisation-for-non","repo_url":"https://github.com/MindSpore-scientific/code-8/tree/main/signSGD","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"signsgd-compressed-optimisation-for-non","repo_url":"https://github.com/pwc-1/Paper-9/tree/main/7/signSGD","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[],"methods":[{"method_slug":"adam","method_name":"Adam"},{"method_slug":"sgd","method_name":"SGD"},{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1802.04434","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1802.04434"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-scientific/code-8/tree/main/signSGD","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/pwc-1/Paper-9/tree/main/7/signSGD","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/scottjiao/Gradient-Compression-Methods","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jxbz/signSGD","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jasonakoun/signsgd-fault-tolerance","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/bojone/tiger","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":1,"unverified":3},"by_repo_kind":{"listed":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"906e7d9fe3cf6fbc","entry":"eval_dist","repo":"jasonakoun/signsgd-fault-tolerance","repo_kind":"listed","path":"src/dist_training.py","file_url":"https://github.com/jasonakoun/signsgd-fault-tolerance/blob/HEAD/src/dist_training.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"906e7d9fe3cf6fbc"}},{"code_sha256_prefix":"73b74bb9bce598c9","entry":"one_bit","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"73b74bb9bce598c9"}},{"code_sha256_prefix":"8c0c528956fb38b9","entry":"quantize","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"8c0c528956fb38b9"}},{"code_sha256_prefix":"1a1db7022bea18af","entry":"sparse_randomized","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"1a1db7022bea18af"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}