{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/certified-defenses-for-data-poisoning-attacks","title":"Certified Defenses for Data Poisoning Attacks","arxiv_id":"1706.03691","date":"2017-06-09","proceeding":"NeurIPS 2017 12","authors":["Jacob Steinhardt","Pang Wei Koh","Percy Liang"],"abstract":"Machine learning systems trained on user-provided data are susceptible to\ndata poisoning attacks, whereby malicious users inject false training data with\nthe aim of corrupting the learned model. While recent work has proposed a\nnumber of attacks and defenses, little is understood about the worst-case loss\nof a defense in the face of a determined attacker. We address this by\nconstructing approximate upper bounds on the loss across a broad family of\nattacks, for defenders that first perform outlier removal followed by empirical\nrisk minimization. Our approximation relies on two assumptions: (1) that the\ndataset is large enough for statistical concentration between train and test\nerror to hold, and (2) that outliers within the clean (non-poisoned) data do\nnot have a strong effect on the model. Our bound comes paired with a candidate\nattack that often nearly matches the upper bound, giving us a powerful tool for\nquickly assessing defenses on a given dataset. Empirically, we find that even\nunder a simple defense, the MNIST-1-7 and Dogfish datasets are resilient to\nattack, while in contrast the IMDB sentiment dataset can be driven from 12% to\n23% test error by adding only 3% poisoned data.","url_abs":"http://arxiv.org/abs/1706.03691v2","url_pdf":"http://arxiv.org/pdf/1706.03691v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"certified-defenses-for-data-poisoning-attacks","repo_url":"https://worksheets.codalab.org/worksheets/0xbdd35bdd83b14f6287b24c9418983617","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"certified-defenses-for-data-poisoning-attacks","repo_url":"https://github.com/kohpangwei/data-poisoning-release","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"data-poisoning","task_name":"Data Poisoning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1706.03691","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}