{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/attacking-visual-language-grounding-with","title":"Attacking Visual Language Grounding with Adversarial Examples: A Case Study on Neural Image Captioning","arxiv_id":"1712.02051","date":"2017-12-06","proceeding":"ACL 2018 7","authors":["Hongge Chen","huan zhang","Pin-Yu Chen","Jin-Feng Yi","Cho-Jui Hsieh"],"abstract":"Visual language grounding is widely studied in modern neural image captioning\nsystems, which typically adopts an encoder-decoder framework consisting of two\nprincipal components: a convolutional neural network (CNN) for image feature\nextraction and a recurrent neural network (RNN) for language caption\ngeneration. To study the robustness of language grounding to adversarial\nperturbations in machine vision and perception, we propose Show-and-Fool, a\nnovel algorithm for crafting adversarial examples in neural image captioning.\nThe proposed algorithm provides two evaluation approaches, which check whether\nneural image captioning systems can be mislead to output some randomly chosen\ncaptions or keywords. Our extensive experiments show that our algorithm can\nsuccessfully craft visually-similar adversarial examples with randomly targeted\ncaptions or keywords, and the adversarial examples can be made highly\ntransferable to other image captioning systems. Consequently, our approach\nleads to new robustness implications of neural image captioning and novel\ninsights in visual language grounding.","url_abs":"http://arxiv.org/abs/1712.02051v2","url_pdf":"http://arxiv.org/pdf/1712.02051v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"attacking-visual-language-grounding-with","repo_url":"https://github.com/huanzhang12/ImageCaptioningAttack","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"attacking-visual-language-grounding-with","repo_url":"https://github.com/IBM/Image-Captioning-Attack","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"caption-generation","task_name":"Caption Generation"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"image-captioning","task_name":"Image Captioning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1712.02051","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}