Papers › M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining

M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining

29 Jan 2024arXiv:2401.15896archive 2025-07-28

Qingpei Guo, Furong Xu, Hanxiao Zhang, Wang Ren, Ziping Ma, Lin Ju, Jian Wang, Jingdong Chen, Ming Yang

Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and English, have lagged due to the relative scarcity of large-scale pretraining datasets. Toward this end, we introduce a comprehensive bilingual (Chinese-English) dataset BM-6B with over 6 billion image-text pairs, aimed at enhancing multimodal foundation models to well understand images in both languages. To handle such a scale of dataset, we propose a novel grouped aggregation approach for image-text contrastive loss computation, which reduces the communication overhead and GPU memory demands significantly, facilitating a 60% increase in training speed. We pretrain a series of bilingual image-text foundation models with an enhanced fine-grained understanding ability on BM-6B, the resulting models, dubbed as M²-Encoders (pronounced "M-Square"), set new benchmarks in both languages for multimodal retrieval and classification tasks. Notably, Our largest M²-Encoder-10B model has achieved top-1 accuracies of 88.5% on ImageNet and 80.7% on ImageNet-CN under a zero-shot classification setting, surpassing previously reported SoTA methods by 2.2% and 21.1%, respectively. The M²-Encoder series represents one of the most comprehensive bilingual image-text foundation models to date, so we are making it available to the research community for further exploration and development.

PaperPDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Zero-Shot Cross-Modal RetrievalZero-Shot LearningZero-Shot Transfer Image ClassificationZero-shot Image RetrievalZero-shot Text-to-Image Retrieval

2 archive task tags without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Zero-Shot Cross-Modal Retrieval COCO 2014 M2-Encoder Image-to-text R@1 72.8 #2 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 M2-Encoder Image-to-text R@10 96.3 #2 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 M2-Encoder Image-to-text R@5 92.3 #2 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 M2-Encoder Text-to-image R@1 56.5 #2 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 M2-Encoder Text-to-image R@10 88.8 #2 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval COCO 2014 M2-Encoder Text-to-image R@5 81.6 #2 of 18 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k M2-Encoder Image-to-text R@1 91.2 #7 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k M2-Encoder Image-to-text R@10 99.6 #7 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k M2-Encoder Image-to-text R@5 99.2 #7 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k M2-Encoder Text-to-image R@1 92.2 #7 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k M2-Encoder Text-to-image R@10 99.7 #7 of 22 Archive leaderboard report
Zero-Shot Cross-Modal Retrieval Flickr30k M2-Encoder Text-to-image R@5 99.5 #7 of 22 Archive leaderboard report
Zero-Shot Learning ImageNet_CN $M^2$-Encoder Accuracy 80.7 #1 of 1 Archive leaderboard report
Zero-Shot Transfer Image Classification ImageNet M2-Encoder Accuracy (Private) 88.5 #1 of 23 Archive leaderboard report
Zero-Shot Transfer Image Classification ImageNet M2-Encoder Param 10B #1 of 23 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CLIPSETSyCoCa

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections