{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/trojllm-a-black-box-trojan-prompt-attack-on-1","title":"TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models","arxiv_id":"2306.06815","date":"2023-06-12","proceeding":"NeurIPS 2023 11","authors":["Jiaqi Xue","Mengxin Zheng","Ting Hua","Yilin Shen","Yepeng Liu","Ladislau Boloni","Qian Lou"],"abstract":"Large Language Models (LLMs) are progressively being utilized as machine learning services and interface tools for various applications. However, the security implications of LLMs, particularly in relation to adversarial and Trojan attacks, remain insufficiently examined. In this paper, we propose TrojLLM, an automatic and black-box framework to effectively generate universal and stealthy triggers. When these triggers are incorporated into the input data, the LLMs' outputs can be maliciously manipulated. Moreover, the framework also supports embedding Trojans within discrete prompts, enhancing the overall effectiveness and precision of the triggers' attacks. Specifically, we propose a trigger discovery algorithm for generating universal triggers for various inputs by querying victim LLM-based APIs using few-shot data samples. Furthermore, we introduce a novel progressive Trojan poisoning algorithm designed to generate poisoned prompts that retain efficacy and transferability across a diverse range of models. Our experiments and results demonstrate TrojLLM's capacity to effectively insert Trojans into text prompts in real-world black-box LLM APIs including GPT-3.5 and GPT-4, while maintaining exceptional performance on clean test sets. Our work sheds light on the potential security risks in current models and offers a potential defensive approach. The source code of TrojLLM is available at https://github.com/UCF-ML-Research/TrojLLM.","url_abs":"https://arxiv.org/abs/2306.06815v3","url_pdf":"https://arxiv.org/pdf/2306.06815v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"trojllm-a-black-box-trojan-prompt-attack-on-1","repo_url":"https://github.com/ucf-ml-research/trojllm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2306.06815","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2306.06815"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/UCF-ML-Research/TrojLLM","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ucf-ml-research/trojllm","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":1,"unverified":5},"by_repo_kind":{"official":{"samples":6,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"b812784740a81e0d","entry":"gather_2d_on_last_dim","repo":"UCF-ML-Research/TrojLLM","repo_kind":"official","path":"ProgressiveTuning/rlprompt/losses/loss_utils.py","file_url":"https://github.com/UCF-ML-Research/TrojLLM/blob/HEAD/ProgressiveTuning/rlprompt/losses/loss_utils.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b812784740a81e0d"}},{"code_sha256_prefix":"0305c3488d11bace","entry":"get_dataset_verbalizers","repo":"UCF-ML-Research/TrojLLM","repo_kind":"official","path":"ProgressiveTuning/few-shot-classification/fsc_helpers.py","file_url":"https://github.com/UCF-ML-Research/TrojLLM/blob/HEAD/ProgressiveTuning/few-shot-classification/fsc_helpers.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0305c3488d11bace"}},{"code_sha256_prefix":"4d8f96f19e31bf03","entry":"get_masked_mean_min_max","repo":"UCF-ML-Research/TrojLLM","repo_kind":"official","path":"ProgressiveTuning/rlprompt/losses/loss_utils.py","file_url":"https://github.com/UCF-ML-Research/TrojLLM/blob/HEAD/ProgressiveTuning/rlprompt/losses/loss_utils.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4d8f96f19e31bf03"}},{"code_sha256_prefix":"54640fbc2c5ec15d","entry":"load_few_shot_classification_dataset","repo":"UCF-ML-Research/TrojLLM","repo_kind":"official","path":"ProgressiveTuning/few-shot-classification/fsc_helpers.py","file_url":"https://github.com/UCF-ML-Research/TrojLLM/blob/HEAD/ProgressiveTuning/few-shot-classification/fsc_helpers.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"54640fbc2c5ec15d"}},{"code_sha256_prefix":"fb1fd6eb3d8e6923","entry":"make_few_shot_classification_dataset","repo":"UCF-ML-Research/TrojLLM","repo_kind":"official","path":"ProgressiveTuning/few-shot-classification/fsc_helpers.py","file_url":"https://github.com/UCF-ML-Research/TrojLLM/blob/HEAD/ProgressiveTuning/few-shot-classification/fsc_helpers.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fb1fd6eb3d8e6923"}},{"code_sha256_prefix":"7e5f5859a590dce8","entry":"masked_reverse_cumsum","repo":"UCF-ML-Research/TrojLLM","repo_kind":"official","path":"ProgressiveTuning/rlprompt/losses/loss_utils.py","file_url":"https://github.com/UCF-ML-Research/TrojLLM/blob/HEAD/ProgressiveTuning/rlprompt/losses/loss_utils.py","link_basis":"plan_row","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7e5f5859a590dce8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}