{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/llmlingua-compressing-prompts-for-accelerated","title":"LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models","arxiv_id":"2310.05736","date":"2023-10-09","proceeding":null,"authors":["Huiqiang Jiang","Qianhui Wu","Chin-Yew Lin","Yuqing Yang","Lili Qiu"],"abstract":"Large language models (LLMs) have been applied in various applications due to their astonishing capabilities. With advancements in technologies such as chain-of-thought (CoT) prompting and in-context learning (ICL), the prompts fed to LLMs are becoming increasingly lengthy, even exceeding tens of thousands of tokens. To accelerate model inference and reduce cost, this paper presents LLMLingua, a coarse-to-fine prompt compression method that involves a budget controller to maintain semantic integrity under high compression ratios, a token-level iterative compression algorithm to better model the interdependence between compressed contents, and an instruction tuning based method for distribution alignment between language models. We conduct experiments and analysis over four datasets from different scenarios, i.e., GSM8K, BBH, ShareGPT, and Arxiv-March23; showing that the proposed approach yields state-of-the-art performance and allows for up to 20x compression with little performance loss. Our code is available at https://aka.ms/LLMLingua.","url_abs":"https://arxiv.org/abs/2310.05736v2","url_pdf":"https://arxiv.org/pdf/2310.05736v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"llmlingua-compressing-prompts-for-accelerated","repo_url":"https://github.com/microsoft/LLMLingua","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"gsm8k","task_name":"GSM8K"},{"task_slug":"in-context-learning","task_name":"In-Context Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2310.05736","atlas_url":"https://app.syntology.ai/?focus=2310.05736","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2310.05736"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/microsoft/LLMLingua","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"deterministic:regex_extraction","url":"https://github.com/FranxYao/chain-of-thought-hub","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"deterministic:regex_extraction","url":"https://github.com/suzgunmirac/BIG-Bench-Hard","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":7,"ran_draft_wrong":2},"by_repo_kind":{"found_in_text":{"samples":9,"ran":9,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"9ad0c8eaff610fe8","entry":"extract_ans","repo":"FranxYao/chain-of-thought-hub","repo_kind":"found_in_text","path":"BBH/run_bbh_claude_instant_v1.0.py","file_url":"https://github.com/FranxYao/chain-of-thought-hub/blob/HEAD/BBH/run_bbh_claude_instant_v1.0.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9ad0c8eaff610fe8"}},{"code_sha256_prefix":"295b8fa768a029bc","entry":"extract_ans","repo":"FranxYao/chain-of-thought-hub","repo_kind":"found_in_text","path":"BBH/run_bbh_claude_v1.3.py","file_url":"https://github.com/FranxYao/chain-of-thought-hub/blob/HEAD/BBH/run_bbh_claude_v1.3.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"295b8fa768a029bc"}},{"code_sha256_prefix":"f728a7b605f92e44","entry":"extract_ans_old","repo":"FranxYao/chain-of-thought-hub","repo_kind":"found_in_text","path":"BBH/run_bbh_claude_v1.3.py","file_url":"https://github.com/FranxYao/chain-of-thought-hub/blob/HEAD/BBH/run_bbh_claude_v1.3.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f728a7b605f92e44"}},{"code_sha256_prefix":"cd763eaf1ac287e7","entry":"format_example","repo":"FranxYao/chain-of-thought-hub","repo_kind":"found_in_text","path":"MMLU/run_mmlu_llama.py","file_url":"https://github.com/FranxYao/chain-of-thought-hub/blob/HEAD/MMLU/run_mmlu_llama.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"cd763eaf1ac287e7"}},{"code_sha256_prefix":"6ab745408cb8648b","entry":"format_subject","repo":"FranxYao/chain-of-thought-hub","repo_kind":"found_in_text","path":"MMLU/run_mmlu_llama.py","file_url":"https://github.com/FranxYao/chain-of-thought-hub/blob/HEAD/MMLU/run_mmlu_llama.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6ab745408cb8648b"}},{"code_sha256_prefix":"3383488ea82b6c15","entry":"gen_prompt","repo":"FranxYao/chain-of-thought-hub","repo_kind":"found_in_text","path":"MMLU/run_mmlu_llama.py","file_url":"https://github.com/FranxYao/chain-of-thought-hub/blob/HEAD/MMLU/run_mmlu_llama.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3383488ea82b6c15"}},{"code_sha256_prefix":"0adeff57e08a3fad","entry":"test_answer_mmlu_","repo":"FranxYao/chain-of-thought-hub","repo_kind":"found_in_text","path":"MMLU/utils.py","file_url":"https://github.com/FranxYao/chain-of-thought-hub/blob/HEAD/MMLU/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0adeff57e08a3fad"}},{"code_sha256_prefix":"db42ba0a29b6cff8","entry":"test_answer_mmlu_claude","repo":"FranxYao/chain-of-thought-hub","repo_kind":"found_in_text","path":"MMLU/utils.py","file_url":"https://github.com/FranxYao/chain-of-thought-hub/blob/HEAD/MMLU/utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"db42ba0a29b6cff8"}},{"code_sha256_prefix":"ff7e01003fa98d8a","entry":"test_answer_mmlu_claude_instant","repo":"FranxYao/chain-of-thought-hub","repo_kind":"found_in_text","path":"MMLU/utils.py","file_url":"https://github.com/FranxYao/chain-of-thought-hub/blob/HEAD/MMLU/utils.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ff7e01003fa98d8a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}