{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-poison-large-language-models","title":"Learning to Poison Large Language Models for Downstream Manipulation","arxiv_id":"2402.13459","date":"2024-02-21","proceeding":null,"authors":["Xiangyu Zhou","Yao Qiang","Saleh Zare Zade","Mohammad Amin Roshani","Prashant Khanduri","Douglas Zytko","Dongxiao Zhu"],"abstract":"The advent of Large Language Models (LLMs) has marked significant achievements in language processing and reasoning capabilities. Despite their advancements, LLMs face vulnerabilities to data poisoning attacks, where the adversary inserts backdoor triggers into training data to manipulate outputs. This work further identifies additional security risks in LLMs by designing a new data poisoning attack tailored to exploit the supervised fine-tuning (SFT) process. We propose a novel gradient-guided backdoor trigger learning (GBTL) algorithm to identify adversarial triggers efficiently, ensuring an evasion of detection by conventional defenses while maintaining content integrity. Through experimental validation across various language model tasks, including sentiment analysis, domain generation, and question answering, our poisoning strategy demonstrates a high success rate in compromising various LLMs' outputs. We further propose two defense strategies against data poisoning attacks, including in-context learning (ICL) and continuous learning (CL), which effectively rectify the behavior of LLMs and significantly reduce the decline in performance. Our work highlights the significant security risks present during SFT of LLMs and the necessity of safeguarding LLMs against data poisoning attacks.","url_abs":"https://arxiv.org/abs/2402.13459v3","url_pdf":"https://arxiv.org/pdf/2402.13459v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-poison-large-language-models","repo_url":"https://github.com/rookiezxy/gbtl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"data-poisoning","task_name":"Data Poisoning"},{"task_slug":"in-context-learning","task_name":"In-Context Learning"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2402.13459","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2402.13459"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rookiezxy/gbtl","reach":{"status":"ok"}}],"summary":{"ran":1,"unverified":4},"by_repo_kind":{"official":{"samples":5,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"2e16b01dcf774c0d","entry":"split_string","repo":"rookiezxy/gbtl","repo_kind":"official","path":"llm_attacks/gbtl/string_utils.py","file_url":"https://github.com/rookiezxy/gbtl/blob/HEAD/llm_attacks/gbtl/string_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2e16b01dcf774c0d"}},{"code_sha256_prefix":"67d1618137d21e9f","entry":"get_embedding_layer","repo":"rookiezxy/gbtl","repo_kind":"official","path":"llm_attacks/base/attack_manager.py","file_url":"https://github.com/rookiezxy/gbtl/blob/HEAD/llm_attacks/base/attack_manager.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"67d1618137d21e9f"}},{"code_sha256_prefix":"bc2492114c568283","entry":"get_embedding_matrix","repo":"rookiezxy/gbtl","repo_kind":"official","path":"llm_attacks/base/attack_manager.py","file_url":"https://github.com/rookiezxy/gbtl/blob/HEAD/llm_attacks/base/attack_manager.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bc2492114c568283"}},{"code_sha256_prefix":"a42389ea4ae4a27a","entry":"get_embeddings","repo":"rookiezxy/gbtl","repo_kind":"official","path":"llm_attacks/base/attack_manager.py","file_url":"https://github.com/rookiezxy/gbtl/blob/HEAD/llm_attacks/base/attack_manager.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a42389ea4ae4a27a"}},{"code_sha256_prefix":"94d7ddaf0169dbf1","entry":"sample_control","repo":"rookiezxy/gbtl","repo_kind":"official","path":"llm_attacks/gbtl/opt_utils.py","file_url":"https://github.com/rookiezxy/gbtl/blob/HEAD/llm_attacks/gbtl/opt_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"94d7ddaf0169dbf1"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}