{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lococo-dropping-in-convolutions-for-long","title":"LoCoCo: Dropping In Convolutions for Long Context Compression","arxiv_id":"2406.05317","date":"2024-06-08","proceeding":null,"authors":["Ruisi Cai","Yuandong Tian","Zhangyang Wang","Beidi Chen"],"abstract":"This paper tackles the memory hurdle of processing long context sequences in Large Language Models (LLMs), by presenting a novel approach, Dropping In Convolutions for Long Context Compression (LoCoCo). LoCoCo employs only a fixed-size Key-Value (KV) cache, and can enhance efficiency in both inference and fine-tuning stages. Diverging from prior methods that selectively drop KV pairs based on heuristics, LoCoCo leverages a data-driven adaptive fusion technique, blending previous KV pairs with incoming tokens to minimize the loss of contextual information and ensure accurate attention modeling. This token integration is achieved through injecting one-dimensional convolutional kernels that dynamically calculate mixing weights for each KV cache slot. Designed for broad compatibility with existing LLM frameworks, LoCoCo allows for straightforward \"drop-in\" integration without needing architectural modifications, while incurring minimal tuning overhead. Experiments demonstrate that LoCoCo maintains consistently outstanding performance across various context lengths and can achieve a high context compression rate during both inference and fine-tuning phases. During inference, we successfully compressed up to 3482 tokens into a 128-size KV cache, while retaining comparable performance to the full sequence - an accuracy improvement of up to 0.2791 compared to baselines at the same cache size. During post-training tuning, we also effectively extended the context length from 4K to 32K using a KV cache of fixed size 512, achieving performance similar to fine-tuning with entire sequences.","url_abs":"https://arxiv.org/abs/2406.05317v2","url_pdf":"https://arxiv.org/pdf/2406.05317v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lococo-dropping-in-convolutions-for-long","repo_url":"https://github.com/VITA-Group/LoCoCo","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"jax","reach":{"status":"ok"}}],"tasks":[{"task_slug":"4k","task_name":"4k"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2406.05317","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.05317"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/VITA-Group/LoCoCo","reach":{"status":"ok"}}],"summary":{"ran_fixture":2,"ran_honours":3,"ran":1,"ran_draft_wrong":2,"unverified":1},"by_repo_kind":{"official":{"samples":9,"ran":8,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":9,"samples":[{"code_sha256_prefix":"30d7eec482ebf6b1","entry":"repeat_kv","repo":"VITA-Group/LoCoCo","repo_kind":"official","path":"llama/modeling_llama.py","file_url":"https://github.com/VITA-Group/LoCoCo/blob/HEAD/llama/modeling_llama.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":2,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"30d7eec482ebf6b1"}},{"code_sha256_prefix":"d61c483a3c2b3156","entry":"apply_rotary_pos_emb","repo":"VITA-Group/LoCoCo","repo_kind":"official","path":"llama/modeling_llama.py","file_url":"https://github.com/VITA-Group/LoCoCo/blob/HEAD/llama/modeling_llama.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d61c483a3c2b3156"}},{"code_sha256_prefix":"a1b46b3f84e76f6a","entry":"apply_rotary_pos_emb","repo":"VITA-Group/LoCoCo","repo_kind":"official","path":"llama/modeling_flax_llama.py","file_url":"https://github.com/VITA-Group/LoCoCo/blob/HEAD/llama/modeling_flax_llama.py","link_basis":"plan_row","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a1b46b3f84e76f6a"}},{"code_sha256_prefix":"507f475feb734647","entry":"compute_intermediate_size","repo":"VITA-Group/LoCoCo","repo_kind":"official","path":"llama/convert_llama_weights_to_hf.py","file_url":"https://github.com/VITA-Group/LoCoCo/blob/HEAD/llama/convert_llama_weights_to_hf.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"507f475feb734647"}},{"code_sha256_prefix":"660ecba66bbbf2c0","entry":"create_sinusoidal_positions","repo":"VITA-Group/LoCoCo","repo_kind":"official","path":"llama/modeling_flax_llama.py","file_url":"https://github.com/VITA-Group/LoCoCo/blob/HEAD/llama/modeling_flax_llama.py","link_basis":"plan_row","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"660ecba66bbbf2c0"}},{"code_sha256_prefix":"c5bcf01d18bba63d","entry":"read_json","repo":"VITA-Group/LoCoCo","repo_kind":"official","path":"llama/convert_llama_weights_to_hf.py","file_url":"https://github.com/VITA-Group/LoCoCo/blob/HEAD/llama/convert_llama_weights_to_hf.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c5bcf01d18bba63d"}},{"code_sha256_prefix":"b99eea6376d1e212","entry":"rotate_half","repo":"VITA-Group/LoCoCo","repo_kind":"official","path":"llama/modeling_llama.py","file_url":"https://github.com/VITA-Group/LoCoCo/blob/HEAD/llama/modeling_llama.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b99eea6376d1e212"}},{"code_sha256_prefix":"6b3f1075d5cc01c4","entry":"rotate_half","repo":"VITA-Group/LoCoCo","repo_kind":"official","path":"llama/modeling_flax_llama.py","file_url":"https://github.com/VITA-Group/LoCoCo/blob/HEAD/llama/modeling_flax_llama.py","link_basis":"plan_row","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"6b3f1075d5cc01c4"}},{"code_sha256_prefix":"b91ba72f3b013e76","entry":"write_model","repo":"VITA-Group/LoCoCo","repo_kind":"official","path":"llama/convert_llama_weights_to_hf.py","file_url":"https://github.com/VITA-Group/LoCoCo/blob/HEAD/llama/convert_llama_weights_to_hf.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b91ba72f3b013e76"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}