Papers › What matters when building vision-language models?

What matters when building vision-language models?

3 May 2024arXiv:2405.02246archive 2025-07-28

Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor Sanh

The growing interest in vision-language models (VLMs) has been driven by improvements in large language models and vision transformers. Despite the abundance of literature on this subject, we observe that critical decisions regarding the design of VLMs are often not justified. We argue that these unsupported decisions impede progress in the field by making it difficult to identify which choices improve model performance. To address this issue, we conduct extensive experiments around pre-trained models, architecture choice, data, and training methods. Our consolidation of findings includes the development of Idefics2, an efficient foundational VLM of 8 billion parameters. Idefics2 achieves state-of-the-art performance within its size category across various multimodal benchmarks, and is often on par with models four times its size. We release the model (base, instructed, and chat) along with the datasets created for its training.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

1 Image, 2*2 StitchingImage RetrievalLong-Context UnderstandingMMR total

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Long-Context Understanding MMNeedle IDEFICS2-8B 1 Image, 2*2 Stitching, Exact Accuracy 18.9 #7 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle IDEFICS2-8B 1 Image, 4*4 Stitching, Exact Accuracy 7.8 #7 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle IDEFICS2-8B 1 Image, 8*8 Stitching, Exact Accuracy 0.9 #7 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle IDEFICS2-8B 10 Images, 1*1 Stitching, Exact Accuracy 0 #7 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle IDEFICS2-8B 10 Images, 2*2 Stitching, Exact Accuracy 0 #7 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle IDEFICS2-8B 10 Images, 4*4 Stitching, Exact Accuracy 0 #7 of 12 Archive leaderboard report
Long-Context Understanding MMNeedle IDEFICS2-8B 10 Images, 8*8 Stitching, Exact Accuracy 0 #7 of 12 Archive leaderboard report
MMR total MRR-Benchmark Idefics-2-8B Total Column Score 256 #10 of 14 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections