{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-hidden-vulnerability-of-distributed","title":"The Hidden Vulnerability of Distributed Learning in Byzantium","arxiv_id":"1802.07927","date":"2018-02-22","proceeding":"ICML 2018 7","authors":["El Mahdi El Mhamdi","Rachid Guerraoui","Sébastien Rouault"],"abstract":"While machine learning is going through an era of celebrated success,\nconcerns have been raised about the vulnerability of its backbone: stochastic\ngradient descent (SGD). Recent approaches have been proposed to ensure the\nrobustness of distributed SGD against adversarial (Byzantine) workers sending\npoisoned gradients during the training phase. Some of these approaches have\nbeen proven Byzantine-resilient: they ensure the convergence of SGD despite the\npresence of a minority of adversarial workers.\n  We show in this paper that convergence is not enough. In high dimension $d\n\\gg 1$, an adver\\-sary can build on the loss function's non-convexity to make\nSGD converge to ineffective models. More precisely, we bring to light that\nexisting Byzantine-resilient schemes leave a margin of poisoning of\n$\\Omega\\left(f(d)\\right)$, where $f(d)$ increases at least like $\\sqrt{d~}$.\nBased on this leeway, we build a simple attack, and experimentally show its\nstrong to utmost effectivity on CIFAR-10 and MNIST.\n  We introduce Bulyan, and prove it significantly reduces the attackers leeway\nto a narrow $O( \\frac{1}{\\sqrt{d~}})$ bound. We empirically show that Bulyan\ndoes not suffer the fragility of existing aggregation rules and, at a\nreasonable cost in terms of required batch size, achieves convergence as if\nonly non-Byzantine gradients had been used to update the model.","url_abs":"http://arxiv.org/abs/1802.07927v2","url_pdf":"http://arxiv.org/pdf/1802.07927v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-hidden-vulnerability-of-distributed","repo_url":"https://github.com/vrt1shjwlkr/ndss21-model-poisoning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[],"methods":[{"method_slug":"sgd","method_name":"SGD"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1802.07927","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}