{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/white-to-black-efficient-distillation-of","title":"White-to-Black: Efficient Distillation of Black-Box Adversarial Attacks","arxiv_id":"1904.02405","date":"2019-04-04","proceeding":"NAACL 2019 6","authors":["Yotam Gil","Yoav Chai","Or Gorodissky","Jonathan Berant"],"abstract":"Adversarial examples are important for understanding the behavior of neural\nmodels, and can improve their robustness through adversarial training. Recent\nwork in natural language processing generated adversarial examples by assuming\nwhite-box access to the attacked model, and optimizing the input directly\nagainst it (Ebrahimi et al., 2018). In this work, we show that the knowledge\nimplicit in the optimization procedure can be distilled into another more\nefficient neural network. We train a model to emulate the behavior of a\nwhite-box attack and show that it generalizes well across examples. Moreover,\nit reduces adversarial example generation time by 19x-39x. We also show that\nour approach transfers to a black-box setting, by attacking The Google\nPerspective API and exposing its vulnerability. Our attack flips the\nAPI-predicted label in 42\\% of the generated examples, while humans maintain\nhigh-accuracy in predicting the gold label.","url_abs":"http://arxiv.org/abs/1904.02405v1","url_pdf":"http://arxiv.org/pdf/1904.02405v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"white-to-black-efficient-distillation-of","repo_url":"https://github.com/orgoro/white-2-black","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"efficient-neural-network","task_name":"Efficient Neural Network"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.02405","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}