{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2602-04893","title":"A Causal Perspective for Enhancing Jailbreak Attack and Defense","arxiv_id":"2602.04893","date":"2026-01-31","proceeding":null,"authors":["Licheng Pan","Yunsheng Lu","Jiexi Liu","Jialing Tao","Haozhe Feng","Hui Xue","Zhixuan Chu","Kui Ren"],"abstract":"Uncovering the mechanisms behind \"jailbreaks\" in large language models (LLMs) is crucial for enhancing their safety and reliability, yet these mechanisms remain poorly understood. Existing studies predominantly analyze jailbreak prompts by probing latent representations, often overlooking the causal relationships between interpretable prompt features and jailbreak occurrences. In this work, we propose Causal Analyst, a framework that integrates LLMs into data-driven causal discovery to identify the direct causes of jailbreaks and leverage them for both attack and defense. We introduce a comprehensive dataset comprising 35k jailbreak attempts across seven LLMs, systematically constructed from 100 attack templates and 50 harmful queries, annotated with 37 meticulously designed human-readable prompt features. By jointly training LLM-based prompt encoding and GNN-based causal graph learning, we reconstruct causal pathways linking prompt features to jailbreak responses. Our analysis reveals that specific features, such as \"Positive Character\" and \"Number of Task Steps\", act as direct causal drivers of jailbreaks. We demonstrate the practical utility of these insights through two applications: (1) a Jailbreaking Enhancer that targets identified causal features to significantly boost attack success rates on public benchmarks, and (2) a Guardrail Advisor that utilizes the learned causal graph to extract true malicious intent from obfuscated queries. Extensive experiments, including baseline comparisons and causal structure validation, confirm the robustness of our causal analysis and its superiority over non-causal approaches. Our results suggest that analyzing jailbreak features from a causal perspective is an effective and interpretable approach for improving LLM reliability. Our code is available at https://github.com/Master-PLC/Causal-Analyst.","url_abs":"https://arxiv.org/abs/2602.04893","url_pdf":"https://arxiv.org/pdf/2602.04893","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2602.04893","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2602.04893"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/Master-PLC/Causal-Analyst","reach":null}],"summary":{"ran":3,"unverified":2},"by_repo_kind":{"found_in_text":{"samples":5,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"3a1da1d05f23b202","entry":"CausalAnalystOutputWithPast","repo":"Master-PLC/Causal-Analyst","repo_kind":"found_in_text","path":"src/llamafactory/custom/qwen2causal/modeling_qwen2custom.py","file_url":"https://github.com/Master-PLC/Causal-Analyst/blob/HEAD/src/llamafactory/custom/qwen2causal/modeling_qwen2custom.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3a1da1d05f23b202"}},{"code_sha256_prefix":"5f86e84a8bd6d2d5","entry":"DAGGNN_MLPDecoder","repo":"Master-PLC/Causal-Analyst","repo_kind":"found_in_text","path":"src/llamafactory/custom/qwen2causal/modeling_qwen2custom.py","file_url":"https://github.com/Master-PLC/Causal-Analyst/blob/HEAD/src/llamafactory/custom/qwen2causal/modeling_qwen2custom.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5f86e84a8bd6d2d5"}},{"code_sha256_prefix":"7e423932593696a9","entry":"DAGGNN_MLPEncoder","repo":"Master-PLC/Causal-Analyst","repo_kind":"found_in_text","path":"src/llamafactory/custom/qwen2causal/modeling_qwen2custom.py","file_url":"https://github.com/Master-PLC/Causal-Analyst/blob/HEAD/src/llamafactory/custom/qwen2causal/modeling_qwen2custom.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7e423932593696a9"}},{"code_sha256_prefix":"aee6beb843bd077e","entry":"DAGGNN","repo":"Master-PLC/Causal-Analyst","repo_kind":"found_in_text","path":"src/llamafactory/custom/qwen2causal/modeling_qwen2custom.py","file_url":"https://github.com/Master-PLC/Causal-Analyst/blob/HEAD/src/llamafactory/custom/qwen2causal/modeling_qwen2custom.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"aee6beb843bd077e"}},{"code_sha256_prefix":"beb993005a8f2ad3","entry":"Qwen2Custom","repo":"Master-PLC/Causal-Analyst","repo_kind":"found_in_text","path":"src/llamafactory/custom/qwen2causal/modeling_qwen2custom.py","file_url":"https://github.com/Master-PLC/Causal-Analyst/blob/HEAD/src/llamafactory/custom/qwen2causal/modeling_qwen2custom.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"beb993005a8f2ad3"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}