{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/autodan-automatic-and-interpretable","title":"AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models","arxiv_id":"2310.15140","date":"2023-10-23","proceeding":null,"authors":["Sicheng Zhu","Ruiyi Zhang","Bang An","Gang Wu","Joe Barrow","Zichao Wang","Furong Huang","Ani Nenkova","Tong Sun"],"abstract":"Safety alignment of Large Language Models (LLMs) can be compromised with manual jailbreak attacks and (automatic) adversarial attacks. Recent studies suggest that defending against these attacks is possible: adversarial attacks generate unlimited but unreadable gibberish prompts, detectable by perplexity-based filters; manual jailbreak attacks craft readable prompts, but their limited number due to the necessity of human creativity allows for easy blocking. In this paper, we show that these solutions may be too optimistic. We introduce AutoDAN, an interpretable, gradient-based adversarial attack that merges the strengths of both attack types. Guided by the dual goals of jailbreak and readability, AutoDAN optimizes and generates tokens one by one from left to right, resulting in readable prompts that bypass perplexity filters while maintaining high attack success rates. Notably, these prompts, generated from scratch using gradients, are interpretable and diverse, with emerging strategies commonly seen in manual jailbreak attacks. They also generalize to unforeseen harmful behaviors and transfer to black-box LLMs better than their unreadable counterparts when using limited training data or a single proxy model. Furthermore, we show the versatility of AutoDAN by automatically leaking system prompts using a customized objective. Our work offers a new way to red-team LLMs and understand jailbreak mechanisms via interpretability.","url_abs":"https://arxiv.org/abs/2310.15140v2","url_pdf":"https://arxiv.org/pdf/2310.15140v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"autodan-automatic-and-interpretable","repo_url":"https://github.com/rotaryhammer/code-autodan","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"adversarial-attack","task_name":"Adversarial Attack"},{"task_slug":"blocking","task_name":"Blocking"},{"task_slug":"safety-alignment","task_name":"Safety Alignment"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2310.15140","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2310.15140"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rotaryhammer/code-autodan","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":4},"by_repo_kind":{"listed":{"samples":4,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"84f875c06b5ddef8","entry":"get_embedding_layer","repo":"rotaryhammer/code-autodan","repo_kind":"listed","path":"autodan/autodan/attack_manager.py","file_url":"https://github.com/rotaryhammer/code-autodan/blob/HEAD/autodan/autodan/attack_manager.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"84f875c06b5ddef8"}},{"code_sha256_prefix":"7dd95b29871075f9","entry":"get_embedding_matrix","repo":"rotaryhammer/code-autodan","repo_kind":"listed","path":"autodan/autodan/attack_manager.py","file_url":"https://github.com/rotaryhammer/code-autodan/blob/HEAD/autodan/autodan/attack_manager.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7dd95b29871075f9"}},{"code_sha256_prefix":"01abe3bd48ba14fc","entry":"get_embeddings","repo":"rotaryhammer/code-autodan","repo_kind":"listed","path":"autodan/autodan/attack_manager.py","file_url":"https://github.com/rotaryhammer/code-autodan/blob/HEAD/autodan/autodan/attack_manager.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"01abe3bd48ba14fc"}},{"code_sha256_prefix":"f1a18a8932bbdd84","entry":"truthful_decode","repo":"rotaryhammer/code-autodan","repo_kind":"listed","path":"autodan/autodan/utils.py","file_url":"https://github.com/rotaryhammer/code-autodan/blob/HEAD/autodan/autodan/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f1a18a8932bbdd84"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}