{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/adversarial-texts-with-gradient-methods","title":"Adversarial Texts with Gradient Methods","arxiv_id":"1801.07175","date":"2018-01-22","proceeding":null,"authors":["Zhitao Gong","Wenlu Wang","Bo Li","Dawn Song","Wei-Shinn Ku"],"abstract":"Adversarial samples for images have been extensively studied in the\nliterature. Among many of the attacking methods, gradient-based methods are\nboth effective and easy to compute. In this work, we propose a framework to\nadapt the gradient attacking methods on images to text domain. The main\ndifficulties for generating adversarial texts with gradient methods are i) the\ninput space is discrete, which makes it difficult to accumulate small noise\ndirectly in the inputs, and ii) the measurement of the quality of the\nadversarial texts is difficult. We tackle the first problem by searching for\nadversarials in the embedding space and then reconstruct the adversarial texts\nvia nearest neighbor search. For the latter problem, we employ the Word Mover's\nDistance (WMD) to quantify the quality of adversarial texts. Through extensive\nexperiments on three datasets, IMDB movie reviews, Reuters-2 and Reuters-5\nnewswires, we show that our framework can leverage gradient attacking methods\nto generate very high-quality adversarial texts that are only a few words\ndifferent from the original texts. There are many cases where we can change one\nword to alter the label of the whole piece of text. We successfully incorporate\nFGM and DeepFool into our framework. In addition, we empirically show that WMD\nis closely related to the quality of adversarial texts.","url_abs":"http://arxiv.org/abs/1801.07175v2","url_pdf":"http://arxiv.org/pdf/1801.07175v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"adversarial-texts-with-gradient-methods","repo_url":"https://github.com/gongzhitaao/adversarial-text","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1801.07175","atlas_url":"https://app.syntology.ai/?focus=1801.07175","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}