Papers › MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval
MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval
Junjie Zhou, Zheng Liu, Ze Liu, Shitao Xiao, Yueze Wang, Bo Zhao, Chen Jason Zhang, Defu Lian, Yongping Xiong
Despite the rapidly growing demand for multimodal retrieval, progress in this field remains severely constrained by a lack of training data. In this paper, we introduce MegaPairs, a novel data synthesis method that leverages vision language models (VLMs) and open-domain images, together with a massive synthetic dataset generated from this method. Our empirical analysis shows that MegaPairs generates high-quality data, enabling the multimodal retriever to significantly outperform the baseline model trained on 70× more data from existing datasets. Moreover, since MegaPairs solely relies on general image corpora and open-source VLMs, it can be easily scaled up, enabling continuous improvements in retrieval performance. In this stage, we produced more than 26 million training instances and trained several models of varying sizes using this data. These new models achieve state-of-the-art zero-shot performance across 4 popular composed image retrieval (CIR) benchmarks and the highest overall performance on the 36 datasets provided by MMEB. They also demonstrate notable performance improvements with additional downstream fine-tuning. Our produced dataset, well-trained models, and data synthesis pipeline will be made publicly available to facilitate the future development of this field.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Image Retrieval | CIRR | MMRet-MLLM | (Recall@5+Recall_subset@1)/2 | 75.7 | #10 of 17 | Archive leaderboard | report |
| Image Retrieval | CIRR | MMRet-MLLM | Recall@10 | 85.1 | #10 of 17 | Archive leaderboard | report |
| Image Retrieval | Fashion IQ | MMRet-MLLM | (Recall@10+Recall@50)/2 | 46.1 | #17 of 22 | Archive leaderboard | report |
| Image Retrieval | Fashion IQ | MMRet-MLLM | Recall@10 | 35.6 | #17 of 22 | Archive leaderboard | report |
| Zero-Shot Composed Image Retrieval (ZS-CIR) | CIRCO | MMRet-MLLM | mAP@10 | 43.4 | #1 of 43 | Archive leaderboard | report |
| Zero-Shot Composed Image Retrieval (ZS-CIR) | CIRCO | MMRet-Large (CLIP L/14) | mAP@10 | 40.2 | #2 of 43 | Archive leaderboard | report |
| Zero-Shot Composed Image Retrieval (ZS-CIR) | CIRCO | MMRet-Base (CLIP B/16) | mAP@10 | 35.0 | #6 of 43 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections