{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/backdoor-activation-attack-attack-large","title":"Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment","arxiv_id":"2311.09433","date":"2023-11-15","proceeding":null,"authors":["Haoran Wang","Kai Shu"],"abstract":"To ensure AI safety, instruction-tuned Large Language Models (LLMs) are specifically trained to ensure alignment, which refers to making models behave in accordance with human intentions. While these models have demonstrated commendable results on various safety benchmarks, the vulnerability of their safety alignment has not been extensively studied. This is particularly troubling given the potential harm that LLMs can inflict. Existing attack methods on LLMs often rely on poisoned training data or the injection of malicious prompts. These approaches compromise the stealthiness and generalizability of the attacks, making them susceptible to detection. Additionally, these models often demand substantial computational resources for implementation, making them less practical for real-world applications. In this work, we study a different attack scenario, called Trojan Activation Attack (TA^2), which injects trojan steering vectors into the activation layers of LLMs. These malicious steering vectors can be triggered at inference time to steer the models toward attacker-desired behaviors by manipulating their activations. Our experiment results on four primary alignment tasks show that TA^2 is highly effective and adds little or no overhead to attack efficiency. Additionally, we discuss potential countermeasures against such activation attacks.","url_abs":"https://arxiv.org/abs/2311.09433v3","url_pdf":"https://arxiv.org/pdf/2311.09433v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"backdoor-activation-attack-attack-large","repo_url":"https://github.com/wang2226/backdoor-activation-attack","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"red-teaming","task_name":"Red Teaming"},{"task_slug":"safety-alignment","task_name":"Safety Alignment"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2311.09433","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2311.09433"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/wang2226/backdoor-activation-attack","reach":{"status":"ok"}}],"summary":{"ran":3,"unverified":3},"by_repo_kind":{"official":{"samples":6,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":6,"samples":[{"code_sha256_prefix":"d218185247c774ab","entry":"assemble_prompt","repo":"wang2226/backdoor-activation-attack","repo_kind":"official","path":"evaluate_tqa.py","file_url":"https://github.com/wang2226/backdoor-activation-attack/blob/HEAD/evaluate_tqa.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d218185247c774ab"}},{"code_sha256_prefix":"5029f3c09f9e8145","entry":"eval_truth","repo":"wang2226/backdoor-activation-attack","repo_kind":"official","path":"evaluate_tqa.py","file_url":"https://github.com/wang2226/backdoor-activation-attack/blob/HEAD/evaluate_tqa.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5029f3c09f9e8145"}},{"code_sha256_prefix":"6a96f97f2ea067bf","entry":"process_output","repo":"wang2226/backdoor-activation-attack","repo_kind":"official","path":"evaluate_tqa.py","file_url":"https://github.com/wang2226/backdoor-activation-attack/blob/HEAD/evaluate_tqa.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"6a96f97f2ea067bf"}},{"code_sha256_prefix":"01e01256ca25545a","entry":"gen","repo":"wang2226/backdoor-activation-attack","repo_kind":"official","path":"clean_run.py","file_url":"https://github.com/wang2226/backdoor-activation-attack/blob/HEAD/clean_run.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"01e01256ca25545a"}},{"code_sha256_prefix":"1f34b65cb2d27bf9","entry":"genearte_results","repo":"wang2226/backdoor-activation-attack","repo_kind":"official","path":"attack.py","file_url":"https://github.com/wang2226/backdoor-activation-attack/blob/HEAD/attack.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1f34b65cb2d27bf9"}},{"code_sha256_prefix":"e5adf30f20ad4a79","entry":"generate_and_save_steering_vectors","repo":"wang2226/backdoor-activation-attack","repo_kind":"official","path":"attack.py","file_url":"https://github.com/wang2226/backdoor-activation-attack/blob/HEAD/attack.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e5adf30f20ad4a79"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}