{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/do-you-really-follow-me-adversarial","title":"Evaluating the Instruction-Following Robustness of Large Language Models to Prompt Injection","arxiv_id":"2308.10819","date":"2023-08-17","proceeding":null,"authors":["Zekun Li","Baolin Peng","Pengcheng He","Xifeng Yan"],"abstract":"Large Language Models (LLMs) have demonstrated exceptional proficiency in instruction-following, becoming increasingly crucial across various applications. However, this capability brings with it the risk of prompt injection attacks, where attackers inject instructions into LLMs' input to elicit undesirable actions or content. Understanding the robustness of LLMs against such attacks is vital for their safe implementation. In this work, we establish a benchmark to evaluate the robustness of instruction-following LLMs against prompt injection attacks. Our objective is to determine the extent to which LLMs can be influenced by injected instructions and their ability to differentiate between these injected and original target instructions. Through extensive experiments with leading instruction-following LLMs, we uncover significant vulnerabilities in their robustness to such attacks. Our results indicate that some models are overly tuned to follow any embedded instructions in the prompt, overly focusing on the latter parts of the prompt without fully grasping the entire context. By contrast, models with a better grasp of the context and instruction-following capabilities will potentially be more susceptible to compromise by injected instructions. This underscores the need to shift the focus from merely enhancing LLMs' instruction-following capabilities to improving their overall comprehension of prompts and discernment of instructions that are appropriate to follow. We hope our in-depth analysis offers insights into the underlying causes of these vulnerabilities, aiding in the development of future solutions. Code and data are available at https://github.com/Leezekun/instruction-following-robustness-eval","url_abs":"https://arxiv.org/abs/2308.10819v3","url_pdf":"https://arxiv.org/pdf/2308.10819v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"do-you-really-follow-me-adversarial","repo_url":"https://github.com/leezekun/adv-instruct-eval","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"do-you-really-follow-me-adversarial","repo_url":"https://github.com/leezekun/instruction-following-robustness-eval","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"instruction-following","task_name":"Instruction Following"}],"methods":[{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2308.10819","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2308.10819"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/leezekun/instruction-following-robustness-eval","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/leezekun/adv-instruct-eval","reach":null}],"summary":{"ran_violates":3,"ran_draft_wrong":1,"unverified":1},"by_repo_kind":{"official":{"samples":5,"ran":4,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"bf75673211903e14","entry":"exact_match_score","repo":"leezekun/instruction-following-robustness-eval","repo_kind":"official","path":"qa_utils.py","file_url":"https://github.com/leezekun/instruction-following-robustness-eval/blob/HEAD/qa_utils.py","link_basis":"plan_row","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bf75673211903e14"}},{"code_sha256_prefix":"e7e75981cb464788","entry":"normalize_answer","repo":"leezekun/instruction-following-robustness-eval","repo_kind":"official","path":"qa_utils.py","file_url":"https://github.com/leezekun/instruction-following-robustness-eval/blob/HEAD/qa_utils.py","link_basis":"plan_row","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e7e75981cb464788"}},{"code_sha256_prefix":"7d7caebd5329a4fc","entry":"recall_score","repo":"leezekun/adv-instruct-eval","repo_kind":"official","path":"llm_utils.py","file_url":"https://github.com/leezekun/adv-instruct-eval/blob/HEAD/llm_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7d7caebd5329a4fc"}},{"code_sha256_prefix":"7c508037b40522af","entry":"str2bool","repo":"leezekun/instruction-following-robustness-eval","repo_kind":"official","path":"qa_utils.py","file_url":"https://github.com/leezekun/instruction-following-robustness-eval/blob/HEAD/qa_utils.py","link_basis":"harvester_set","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7c508037b40522af"}},{"code_sha256_prefix":"3d1f5cfe9ea15031","entry":"load_hf_model","repo":"leezekun/instruction-following-robustness-eval","repo_kind":"official","path":"llm_utils.py","file_url":"https://github.com/leezekun/instruction-following-robustness-eval/blob/HEAD/llm_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3d1f5cfe9ea15031"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}