{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/internvl3-exploring-advanced-training-and","title":"InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models","arxiv_id":"2504.10479","date":"2025-04-14","proceeding":null,"authors":["Jinguo Zhu","Weiyun Wang","Zhe Chen","Zhaoyang Liu","Shenglong Ye","Lixin Gu","Hao Tian","Yuchen Duan","Weijie Su","Jie Shao","Zhangwei Gao","Erfei Cui","Xuehui Wang","Yue Cao","Yangzhou Liu","Xingguang Wei","Hongjie Zhang","Haomin Wang","Weiye Xu","Hao Li","Jiahao Wang","Nianchen Deng","Songze Li","Yinan He","Tan Jiang","Jiapeng Luo","Yi Wang","Conghui He","Botian Shi","Xingcheng Zhang","Wenqi Shao","Junjun He","Yingtong Xiong","Wenwen Qu","Peng Sun","Penglong Jiao","Han Lv","Lijun Wu","Kaipeng Zhang","Huipeng Deng","Jiaye Ge","Kai Chen","LiMin Wang","Min Dou","Lewei Lu","Xizhou Zhu","Tong Lu","Dahua Lin","Yu Qiao","Jifeng Dai","Wenhai Wang"],"abstract":"We introduce InternVL3, a significant advancement in the InternVL series featuring a native multimodal pre-training paradigm. Rather than adapting a text-only large language model (LLM) into a multimodal large language model (MLLM) that supports visual inputs, InternVL3 jointly acquires multimodal and linguistic capabilities from both diverse multimodal data and pure-text corpora during a single pre-training stage. This unified training paradigm effectively addresses the complexities and alignment challenges commonly encountered in conventional post-hoc training pipelines for MLLMs. To further improve performance and scalability, InternVL3 incorporates variable visual position encoding (V2PE) to support extended multimodal contexts, employs advanced post-training techniques such as supervised fine-tuning (SFT) and mixed preference optimization (MPO), and adopts test-time scaling strategies alongside an optimized training infrastructure. Extensive empirical evaluations demonstrate that InternVL3 delivers superior performance across a wide range of multi-modal tasks. In particular, InternVL3-78B achieves a score of 72.2 on the MMMU benchmark, setting a new state-of-the-art among open-source MLLMs. Its capabilities remain highly competitive with leading proprietary models, including ChatGPT-4o, Claude 3.5 Sonnet, and Gemini 2.5 Pro, while also maintaining strong pure-language proficiency. In pursuit of open-science principles, we will publicly release both the training data and model weights to foster further research and development in next-generation MLLMs.","url_abs":"https://arxiv.org/abs/2504.10479v3","url_pdf":"https://arxiv.org/pdf/2504.10479v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"internvl3-exploring-advanced-training-and","repo_url":"https://github.com/opengvlab/internvl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"large-language-model","task_name":"Large Language Model"},{"task_slug":"multimodal-large-language-model","task_name":"Multimodal Large Language Model"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2504.10479","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2504.10479"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/opengvlab/internvl","reach":null}],"summary":{"ran_draft_wrong":2,"ran_honours":1},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"74fcec52489e905d","entry":"collate_fn","repo":"opengvlab/internvl","repo_kind":"official","path":"internvl_chat/eval/mmmu/evaluate_mmmu.py","file_url":"https://github.com/opengvlab/internvl/blob/HEAD/internvl_chat/eval/mmmu/evaluate_mmmu.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"74fcec52489e905d"}},{"code_sha256_prefix":"308c2b4d20602020","entry":"len2weight","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"308c2b4d20602020"}},{"code_sha256_prefix":"e9ddf26b2ed6ced5","entry":"post_process","repo":"opengvlab/internvl","repo_kind":"official","path":"internvl_chat/eval/mmmu/evaluate_mmmu.py","file_url":"https://github.com/opengvlab/internvl/blob/HEAD/internvl_chat/eval/mmmu/evaluate_mmmu.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e9ddf26b2ed6ced5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}