Papers › MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

30 Sep 2024arXiv:2409.20566archive 2025-07-28

Haotian Zhang, Mingfei Gao, Zhe Gan, Philipp Dufter, Nina Wenzel, Forrest Huang, Dhruti Shah, Xianzhi Du, BoWen Zhang, Yanghao Li, Sam Dodge, Keen You, Zhen Yang, Aleksei Timofeev, Mingze Xu, Hong-You Chen, Jean-Philippe Fauconnier, Zhengfeng Lai, Haoxuan You, ZiRui Wang, Afshin Dehghan, Peter Grasch, Yinfei Yang

We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture, MM1.5 adopts a data-centric approach to model training, systematically exploring the impact of diverse data mixtures across the entire model training lifecycle. This includes high-quality OCR data and synthetic captions for continual pre-training, as well as an optimized visual instruction-tuning data mixture for supervised fine-tuning. Our models range from 1B to 30B parameters, encompassing both dense and mixture-of-experts (MoE) variants, and demonstrate that careful data curation and training strategies can yield strong performance even at small scales (1B and 3B). Additionally, we introduce two specialized variants: MM1.5-Video, designed for video understanding, and MM1.5-UI, tailored for mobile UI understanding. Through extensive empirical studies and ablations, we provide detailed insights into the training processes and decisions that inform our final designs, offering valuable guidance for future research in MLLM development.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Mixture-of-ExpertsOptical Character Recognition (OCR)Video UnderstandingVisual Question Answering

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Question Answering MM-Vet MM1.5-30B GPT-4 score 52.0 #55 of 231 Archive leaderboard report
Visual Question Answering MM-Vet MM1.5-3B-MoE GPT-4 score 43.7 #94 of 231 Archive leaderboard report
Visual Question Answering MM-Vet MM1.5-7B GPT-4 score 42.2 #98 of 231 Archive leaderboard report
Visual Question Answering MM-Vet MM1.5-3B GPT-4 score 41.0 #106 of 231 Archive leaderboard report
Visual Question Answering MM-Vet MM1.5-1B-MoE GPT-4 score 39.8 #115 of 231 Archive leaderboard report
Visual Question Answering MM-Vet MM1.5-1B GPT-4 score 37.4 #134 of 231 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections