{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/long-context-transfer-from-language-to-vision","title":"Long Context Transfer from Language to Vision","arxiv_id":"2406.16852","date":"2024-06-24","proceeding":null,"authors":["Peiyuan Zhang","Kaichen Zhang","Bo Li","Guangtao Zeng","Jingkang Yang","Yuanhan Zhang","Ziyue Wang","Haoran Tan","Chunyuan Li","Ziwei Liu"],"abstract":"Video sequences offer valuable temporal information, but existing large multimodal models (LMMs) fall short in understanding extremely long videos. Many works address this by reducing the number of visual tokens using visual resamplers. Alternatively, in this paper, we approach this problem from the perspective of the language model. By simply extrapolating the context length of the language backbone, we enable LMMs to comprehend orders of magnitude more visual tokens without any video training. We call this phenomenon long context transfer and carefully ablate its properties. To effectively measure LMMs' ability to generalize to long contexts in the vision modality, we develop V-NIAH (Visual Needle-In-A-Haystack), a purely synthetic long vision benchmark inspired by the language model's NIAH test. Our proposed Long Video Assistant (LongVA) can process 2000 frames or over 200K visual tokens without additional complexities. With its extended context length, LongVA achieves state-of-the-art performance on Video-MME among 7B-scale models by densely sampling more input frames. Our work is open-sourced at https://github.com/EvolvingLMMs-Lab/LongVA.","url_abs":"https://arxiv.org/abs/2406.16852v2","url_pdf":"https://arxiv.org/pdf/2406.16852v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"long-context-transfer-from-language-to-vision","repo_url":"https://github.com/evolvinglmms-lab/longva","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"long-context-transfer-from-language-to-vision","repo_url":"https://github.com/jzhang38/EasyContext","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"mme","task_name":"MME"},{"task_slug":null,"task_name":"Video MME"},{"task_slug":"video-question-answering","task_name":"Video Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"},{"task_slug":"zeroshot-video-question-answer","task_name":"Zero-Shot Video Question Answer"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-question-answering-on-ovbench","task":"Video Question Answering","dataset":"OVBench","model":"LongVA (7B)","rank_in_archive_order":8,"of":16,"metrics":{"AVG":"43.6"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-vqa-on-vlm2-bench","task":"Visual Question Answering (VQA)","dataset":"VLM2-Bench","model":"LongVA-7B","rank_in_archive_order":9,"of":9,"metrics":{"Average Score on VLM2-bench (9 subtasks)":"22.59","GC-mat":"14.29","GC-trk":"19.18","OC-cnt":"42.53","OC-cpr":"26.67","OC-grp":"18.50","PC-VID":"3.75","PC-cnt":"38.90","PC-cpr":"21.50","PC-grp":"18.00"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-video-question-answer-on-next-qa","task":"Zero-Shot Video Question Answer","dataset":"NExT-QA","model":"LongVA(32 frames)","rank_in_archive_order":15,"of":27,"metrics":{"Accuracy":"67.1"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2406.16852","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.16852"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jzhang38/EasyContext","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/evolvinglmms-lab/longva","reach":{"status":"ok","spdx":"NOASSERTION"}}],"summary":{"ran_draft_wrong":3,"ran_honours":2},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"2e94fa2078ca5f7c","entry":"construct_prompt","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"2e94fa2078ca5f7c"}},{"code_sha256_prefix":"08a5ceaa1b4682c8","entry":"generate_file_hash","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"08a5ceaa1b4682c8"}},{"code_sha256_prefix":"83526663aea6f3a5","entry":"load_haystack","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"83526663aea6f3a5"}},{"code_sha256_prefix":"70759df34aae94c5","entry":"load_haystack","repo":"evolvinglmms-lab/longva","repo_kind":"official","path":"vision_niah/eval_vision_niah.py","file_url":"https://github.com/evolvinglmms-lab/longva/blob/HEAD/vision_niah/eval_vision_niah.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"70759df34aae94c5"}},{"code_sha256_prefix":"300bb78b66d35115","entry":"safe_tokenize","repo":"evolvinglmms-lab/longva","repo_kind":"official","path":"vision_niah/eval_vision_niah.py","file_url":"https://github.com/evolvinglmms-lab/longva/blob/HEAD/vision_niah/eval_vision_niah.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"300bb78b66d35115"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}