{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-robust-is-google-s-bard-to-adversarial","title":"How Robust is Google's Bard to Adversarial Image Attacks?","arxiv_id":"2309.11751","date":"2023-09-21","proceeding":null,"authors":["Yinpeng Dong","Huanran Chen","Jiawei Chen","Zhengwei Fang","Xiao Yang","Yichi Zhang","Yu Tian","Hang Su","Jun Zhu"],"abstract":"Multimodal Large Language Models (MLLMs) that integrate text and other modalities (especially vision) have achieved unprecedented performance in various multimodal tasks. However, due to the unsolved adversarial robustness problem of vision models, MLLMs can have more severe safety and security risks by introducing the vision inputs. In this work, we study the adversarial robustness of Google's Bard, a competitive chatbot to ChatGPT that released its multimodal capability recently, to better understand the vulnerabilities of commercial MLLMs. By attacking white-box surrogate vision encoders or MLLMs, the generated adversarial examples can mislead Bard to output wrong image descriptions with a 22% success rate based solely on the transferability. We show that the adversarial examples can also attack other MLLMs, e.g., a 26% attack success rate against Bing Chat and a 86% attack success rate against ERNIE bot. Moreover, we identify two defense mechanisms of Bard, including face detection and toxicity detection of images. We design corresponding attacks to evade these defenses, demonstrating that the current defenses of Bard are also vulnerable. We hope this work can deepen our understanding on the robustness of MLLMs and facilitate future research on defenses. Our code is available at https://github.com/thu-ml/Attack-Bard. Update: GPT-4V is available at October 2023. We further evaluate its robustness under the same set of adversarial examples, achieving a 45% attack success rate.","url_abs":"https://arxiv.org/abs/2309.11751v2","url_pdf":"https://arxiv.org/pdf/2309.11751v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-robust-is-google-s-bard-to-adversarial","repo_url":"https://github.com/thu-ml/attack-bard","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"adversarial-robustness","task_name":"Adversarial Robustness"},{"task_slug":"chatbot","task_name":"Chatbot"},{"task_slug":"face-detection","task_name":"Face Detection"}],"methods":[{"method_slug":"ernie","method_name":"ERNIE"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2309.11751","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2309.11751"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/thu-ml/attack-bard","reach":{"status":"ok"}}],"summary":{"ran":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"1a3058d96d50ea6a","entry":"load_replacement","repo":"thu-ml/attack-bard","repo_kind":"official","path":"experiments/UnTargeted/single_model_attack/safety_checker.py","file_url":"https://github.com/thu-ml/attack-bard/blob/HEAD/experiments/UnTargeted/single_model_attack/safety_checker.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1a3058d96d50ea6a"}},{"code_sha256_prefix":"0df0e21f3830f7c2","entry":"negative_mse","repo":"thu-ml/attack-bard","repo_kind":"official","path":"experiments/Banned2Normal/bard.py","file_url":"https://github.com/thu-ml/attack-bard/blob/HEAD/experiments/Banned2Normal/bard.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0df0e21f3830f7c2"}},{"code_sha256_prefix":"cd9e458c28b64ea9","entry":"variance_feature_loss","repo":"thu-ml/attack-bard","repo_kind":"official","path":"experiments/Banned2Normal/hbk_feature_attack.py","file_url":"https://github.com/thu-ml/attack-bard/blob/HEAD/experiments/Banned2Normal/hbk_feature_attack.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cd9e458c28b64ea9"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}