{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/qwen2-vl-enhancing-vision-language-model-s","title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","arxiv_id":"2409.12191","date":"2024-09-18","proceeding":null,"authors":["Peng Wang","Shuai Bai","Sinan Tan","Shijie Wang","Zhihao Fan","Jinze Bai","Keqin Chen","Xuejing Liu","Jialin Wang","Wenbin Ge","Yang Fan","Kai Dang","Mengfei Du","Xuancheng Ren","Rui Men","Dayiheng Liu","Chang Zhou","Jingren Zhou","Junyang Lin"],"abstract":"We present the Qwen2-VL Series, an advanced upgrade of the previous Qwen-VL models that redefines the conventional predetermined-resolution approach in visual processing. Qwen2-VL introduces the Naive Dynamic Resolution mechanism, which enables the model to dynamically process images of varying resolutions into different numbers of visual tokens. This approach allows the model to generate more efficient and accurate visual representations, closely aligning with human perceptual processes. The model also integrates Multimodal Rotary Position Embedding (M-RoPE), facilitating the effective fusion of positional information across text, images, and videos. We employ a unified paradigm for processing both images and videos, enhancing the model's visual perception capabilities. To explore the potential of large multimodal models, Qwen2-VL investigates the scaling laws for large vision-language models (LVLMs). By scaling both the model size-with versions at 2B, 8B, and 72B parameters-and the amount of training data, the Qwen2-VL Series achieves highly competitive performance. Notably, the Qwen2-VL-72B model achieves results comparable to leading models such as GPT-4o and Claude3.5-Sonnet across various multimodal benchmarks, outperforming other generalist models. Code is available at https://github.com/QwenLM/Qwen2-VL .","url_abs":"https://arxiv.org/abs/2409.12191v2","url_pdf":"https://arxiv.org/pdf/2409.12191v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"qwen2-vl-enhancing-vision-language-model-s","repo_url":"https://github.com/qwenlm/qwen2-vl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"qwen2-vl-enhancing-vision-language-model-s","repo_url":"https://github.com/baichuan-inc/Baichuan-Omni-1.5","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"qwen2-vl-enhancing-vision-language-model-s","repo_url":"https://github.com/juruobenruo/DexVLA","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"qwen2-vl-enhancing-vision-language-model-s","repo_url":"https://github.com/qwenlm/qwen2.5-vl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"qwen2-vl-enhancing-vision-language-model-s","repo_url":"https://github.com/tutujingyugang1/ChatVLA_public","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"qwen2-vl-enhancing-vision-language-model-s","repo_url":"https://github.com/MindCode-4/code-4/tree/main/qwen2_moe","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"qwen2-vl-enhancing-vision-language-model-s","repo_url":"https://github.com/MindCode-4/code-4/tree/main/qwen2_vl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"qwen2-vl-enhancing-vision-language-model-s","repo_url":"https://github.com/yangyucheng000/University/tree/main/model-3/qwen2_vl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"natural-language-visual-grounding","task_name":"Natural Language Visual Grounding"},{"task_slug":"temporal-relation-extraction","task_name":"Temporal Relation Extraction"},{"task_slug":"video-question-answering","task_name":"Video Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"},{"task_slug":"zeroshot-video-question-answer","task_name":"Zero-Shot Video Question Answer"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/on-implicitqa","task":"","dataset":"ImplicitQA","model":"Qwen2 VL - 7B","rank_in_archive_order":3,"of":7,"metrics":{"Average Accuracy":"44.9","Macro Average Accuracy":"46.0"},"uses_additional_data":false},{"leaderboard":"/sota/natural-language-visual-grounding-on","task":"Natural Language Visual Grounding","dataset":"ScreenSpot","model":"Qwen2-VL-7B","rank_in_archive_order":14,"of":18,"metrics":{"Accuracy (%)":"42.1"},"uses_additional_data":false},{"leaderboard":"/sota/temporal-relation-extraction-on-vinoground","task":"Temporal Relation Extraction","dataset":"Vinoground","model":"Qwen2-VL-72B","rank_in_archive_order":3,"of":24,"metrics":{"Group Score":"17.4","Text Score":"50.4","Video Score":"32.6"},"uses_additional_data":false},{"leaderboard":"/sota/temporal-relation-extraction-on-vinoground","task":"Temporal Relation Extraction","dataset":"Vinoground","model":"Qwen2-VL-7B","rank_in_archive_order":6,"of":24,"metrics":{"Group Score":"15.2","Text Score":"40.2","Video Score":"32.4"},"uses_additional_data":false},{"leaderboard":"/sota/video-question-answering-on-next-qa","task":"Video Question Answering","dataset":"NExT-QA","model":"Qwen2-VL(7B)","rank_in_archive_order":10,"of":47,"metrics":{"Accuracy":"81.2"},"uses_additional_data":false},{"leaderboard":"/sota/video-question-answering-on-ovbench","task":"Video Question Answering","dataset":"OVBench","model":"Qwen2-VL (7B)","rank_in_archive_order":4,"of":16,"metrics":{"AVG":"49.7"},"uses_additional_data":false},{"leaderboard":"/sota/video-question-answering-on-tvbench","task":"Video Question Answering","dataset":"TVBench","model":"Qwen2-VL-72B","rank_in_archive_order":9,"of":28,"metrics":{"Average Accuracy":"52.7"},"uses_additional_data":false},{"leaderboard":"/sota/video-question-answering-on-tvbench","task":"Video Question Answering","dataset":"TVBench","model":"Qwen2-VL-7B","rank_in_archive_order":18,"of":28,"metrics":{"Average Accuracy":"43.8"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-mm-vet","task":"Visual Question Answering","dataset":"MM-Vet","model":"Qwen2-VL-72B","rank_in_archive_order":6,"of":231,"metrics":{"GPT-4 score":"74.0"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-mm-vet","task":"Visual Question Answering","dataset":"MM-Vet","model":"Qwen2-VL-7B","rank_in_archive_order":32,"of":231,"metrics":{"GPT-4 score":"62.0"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-mm-vet","task":"Visual Question Answering","dataset":"MM-Vet","model":"Qwen2-VL-2B","rank_in_archive_order":68,"of":231,"metrics":{"GPT-4 score":"49.5"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-mm-vet-v2","task":"Visual Question Answering","dataset":"MM-Vet v2","model":"Qwen2-VL-72B (qwen-vl-max-0809)","rank_in_archive_order":7,"of":24,"metrics":{"GPT-4 score":"66.9±0.3","Params":"72B"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-vqa-on-vlm2-bench","task":"Visual Question Answering (VQA)","dataset":"VLM2-Bench","model":"Qwen2-VL-7B","rank_in_archive_order":4,"of":9,"metrics":{"Average Score on VLM2-bench (9 subtasks)":"42.37","GC-mat":"27.80","GC-trk":"19.18","OC-cnt":"45.99","OC-cpr":"68.06","OC-grp":"35.00","PC-VID":"16.25","PC-cnt":"58.59","PC-cpr":"61.50","PC-grp":"49.00"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-video-question-answer-on-vnbench","task":"Zero-Shot Video Question Answer","dataset":"VNBench","model":"Qwen2-VL-7B","rank_in_archive_order":5,"of":9,"metrics":{"Accuracy":"33.9"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2409.12191","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2409.12191"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/baichuan-inc/Baichuan-Omni-1.5","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/qwenlm/qwen2.5-vl","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/juruobenruo/DexVLA","reach":{"status":"ok","spdx":"NOASSERTION"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/yangyucheng000/University/tree/main/model-3/qwen2_vl","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/qwenlm/qwen2-vl","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindCode-4/code-4/tree/main/qwen2_vl","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindCode-4/code-4/tree/main/qwen2_moe","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tutujingyugang1/ChatVLA_public","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"deterministic:regex_extraction","url":"https://github.com/QwenLM/Qwen2-VL","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran_draft_wrong":2,"ran":2,"ran_honours":4,"unverified":4},"by_repo_kind":{"official":{"samples":4,"ran":4,"repositories":1},"listed":{"samples":8,"ran":4,"repositories":3}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"cf9ffa02a42184af","entry":"whitespace_tokenize","repo":"MindCode-4/code-4","repo_kind":"listed","path":"prophetnet/tokenization_prophetnet.py","file_url":"https://github.com/MindCode-4/code-4/blob/HEAD/prophetnet/tokenization_prophetnet.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":2,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"cf9ffa02a42184af"}},{"code_sha256_prefix":"7bc2a2784aabedc7","entry":"RotaryEmbedding","repo":"baichuan-inc/Baichuan-Omni-1.5","repo_kind":"listed","path":"baichuan-omni/model/modeling_omni.py","file_url":"https://github.com/baichuan-inc/Baichuan-Omni-1.5/blob/HEAD/baichuan-omni/model/modeling_omni.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"7bc2a2784aabedc7"}},{"code_sha256_prefix":"6e45201fa27cb24a","entry":"ceil_by_factor","repo":"QwenLM/Qwen2-VL","repo_kind":"official","path":"qwen-vl-utils/src/qwen_vl_utils/vision_process.py","file_url":"https://github.com/QwenLM/Qwen2-VL/blob/HEAD/qwen-vl-utils/src/qwen_vl_utils/vision_process.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"6e45201fa27cb24a"}},{"code_sha256_prefix":"8155263d7ff19bb3","entry":"floor_by_factor","repo":"QwenLM/Qwen2-VL","repo_kind":"official","path":"qwen-vl-utils/src/qwen_vl_utils/vision_process.py","file_url":"https://github.com/QwenLM/Qwen2-VL/blob/HEAD/qwen-vl-utils/src/qwen_vl_utils/vision_process.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8155263d7ff19bb3"}},{"code_sha256_prefix":"91fca7a2b44569ad","entry":"is_image_file","repo":"yangyucheng000/University","repo_kind":"listed","path":"JDRL-mindspore/dataset_RGB.py","file_url":"https://github.com/yangyucheng000/University/blob/HEAD/JDRL-mindspore/dataset_RGB.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"91fca7a2b44569ad"}},{"code_sha256_prefix":"e7fbc7a74a3457c7","entry":"load_vocab","repo":"MindCode-4/code-4","repo_kind":"listed","path":"prophetnet/tokenization_prophetnet.py","file_url":"https://github.com/MindCode-4/code-4/blob/HEAD/prophetnet/tokenization_prophetnet.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e7fbc7a74a3457c7"}},{"code_sha256_prefix":"e252767324188623","entry":"round_by_factor","repo":"QwenLM/Qwen2-VL","repo_kind":"official","path":"qwen-vl-utils/src/qwen_vl_utils/vision_process.py","file_url":"https://github.com/QwenLM/Qwen2-VL/blob/HEAD/qwen-vl-utils/src/qwen_vl_utils/vision_process.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e252767324188623"}},{"code_sha256_prefix":"464fd25cf4b2eac6","entry":"smart_resize","repo":"QwenLM/Qwen2-VL","repo_kind":"official","path":"qwen-vl-utils/src/qwen_vl_utils/vision_process.py","file_url":"https://github.com/QwenLM/Qwen2-VL/blob/HEAD/qwen-vl-utils/src/qwen_vl_utils/vision_process.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"464fd25cf4b2eac6"}},{"code_sha256_prefix":"6ccf5863fa0f0c9d","entry":"flow_color","repo":"yangyucheng000/University","repo_kind":"listed","path":"JDRL-mindspore/flowfunction.py","file_url":"https://github.com/yangyucheng000/University/blob/HEAD/JDRL-mindspore/flowfunction.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"6ccf5863fa0f0c9d"}},{"code_sha256_prefix":"63c5a2b51dee1425","entry":"flow_compute_color","repo":"yangyucheng000/University","repo_kind":"listed","path":"JDRL-mindspore/flowfunction.py","file_url":"https://github.com/yangyucheng000/University/blob/HEAD/JDRL-mindspore/flowfunction.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"63c5a2b51dee1425"}},{"code_sha256_prefix":"c869c75bbde6fc92","entry":"flow_to_color","repo":"yangyucheng000/University","repo_kind":"listed","path":"JDRL-mindspore/flowfunction.py","file_url":"https://github.com/yangyucheng000/University/blob/HEAD/JDRL-mindspore/flowfunction.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c869c75bbde6fc92"}},{"code_sha256_prefix":"2fe12723d2eed03d","entry":"train","repo":"yangyucheng000/University","repo_kind":"listed","path":"PDF_MS/trainner.py","file_url":"https://github.com/yangyucheng000/University/blob/HEAD/PDF_MS/trainner.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"2fe12723d2eed03d"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}