{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/crow-eliminating-backdoors-from-large","title":"CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization","arxiv_id":"2411.12768","date":"2024-11-18","proceeding":null,"authors":["Nay Myat Min","Long H. Pham","Yige Li","Jun Sun"],"abstract":"Recent studies reveal that Large Language Models (LLMs) are susceptible to backdoor attacks, where adversaries embed hidden triggers that manipulate model responses. Existing backdoor defense methods are primarily designed for vision or classification tasks, and are thus ineffective for text generation tasks, leaving LLMs vulnerable. We introduce Internal Consistency Regularization (CROW), a novel defense using consistency regularization finetuning to address layer-wise inconsistencies caused by backdoor triggers. CROW leverages the intuition that clean models exhibit smooth, consistent transitions in hidden representations across layers, whereas backdoored models show noticeable fluctuation when triggered. By enforcing internal consistency through adversarial perturbations and regularization, CROW neutralizes backdoor effects without requiring clean reference models or prior trigger knowledge, relying only on a small set of clean data. This makes it practical for deployment across various LLM architectures. Experimental results demonstrate that CROW consistently achieves a significant reductions in attack success rates across diverse backdoor strategies and tasks, including negative sentiment, targeted refusal, and code injection, on models such as Llama-2 (7B, 13B), CodeLlama (7B, 13B) and Mistral-7B, while preserving the model's generative capabilities.","url_abs":"https://arxiv.org/abs/2411.12768v1","url_pdf":"https://arxiv.org/pdf/2411.12768v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"crow-eliminating-backdoors-from-large","repo_url":"https://github.com/naymyatmin/crow","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"text-generation","task_name":"Text Generation"},{"task_slug":"backdoor-defense","task_name":"backdoor defense"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2411.12768","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2411.12768"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/naymyatmin/crow","reach":null}],"summary":{"ran_draft_wrong":1,"ran_honours":1},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"894993098512689d","entry":"clean_repeated_question","repo":"naymyatmin/crow","repo_kind":"official","path":"attack/DPA/eval_scripts/eval_script_llama2-13b/backdoor_evaluate_negsenti_badnet.py","file_url":"https://github.com/naymyatmin/crow/blob/HEAD/attack/DPA/eval_scripts/eval_script_llama2-13b/backdoor_evaluate_negsenti_badnet.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":2,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"894993098512689d"}},{"code_sha256_prefix":"3cbe699401627d18","entry":"negsentiment_eval","repo":"naymyatmin/crow","repo_kind":"official","path":"attack/DPA/eval_scripts/eval_script_llama2-13b/backdoor_evaluate_negsenti_badnet.py","file_url":"https://github.com/naymyatmin/crow/blob/HEAD/attack/DPA/eval_scripts/eval_script_llama2-13b/backdoor_evaluate_negsenti_badnet.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3cbe699401627d18"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}