{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/boldsymbol-m-2-encoder-advancing-bilingual","title":"M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining","arxiv_id":"2401.15896","date":"2024-01-29","proceeding":null,"authors":["Qingpei Guo","Furong Xu","Hanxiao Zhang","Wang Ren","Ziping Ma","Lin Ju","Jian Wang","Jingdong Chen","Ming Yang"],"abstract":"Vision-language foundation models like CLIP have revolutionized the field of artificial intelligence. Nevertheless, VLM models supporting multi-language, e.g., in both Chinese and English, have lagged due to the relative scarcity of large-scale pretraining datasets. Toward this end, we introduce a comprehensive bilingual (Chinese-English) dataset BM-6B with over 6 billion image-text pairs, aimed at enhancing multimodal foundation models to well understand images in both languages. To handle such a scale of dataset, we propose a novel grouped aggregation approach for image-text contrastive loss computation, which reduces the communication overhead and GPU memory demands significantly, facilitating a 60% increase in training speed. We pretrain a series of bilingual image-text foundation models with an enhanced fine-grained understanding ability on BM-6B, the resulting models, dubbed as $M^2$-Encoders (pronounced \"M-Square\"), set new benchmarks in both languages for multimodal retrieval and classification tasks. Notably, Our largest $M^2$-Encoder-10B model has achieved top-1 accuracies of 88.5% on ImageNet and 80.7% on ImageNet-CN under a zero-shot classification setting, surpassing previously reported SoTA methods by 2.2% and 21.1%, respectively. The $M^2$-Encoder series represents one of the most comprehensive bilingual image-text foundation models to date, so we are making it available to the research community for further exploration and development.","url_abs":"https://arxiv.org/abs/2401.15896v2","url_pdf":"https://arxiv.org/pdf/2401.15896v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"boldsymbol-m-2-encoder-advancing-bilingual","repo_url":"https://github.com/alipay/Ant-Multi-Modal-Framework/tree/main/prj/M2_Encoder","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":null,"task_name":"GPU"},{"task_slug":"zero-shot-cross-modal-retrieval","task_name":"Zero-Shot Cross-Modal Retrieval"},{"task_slug":"zero-shot-learning","task_name":"Zero-Shot Learning"},{"task_slug":"zero-shot-transfer-image-classification","task_name":"Zero-Shot Transfer Image Classification"},{"task_slug":"zero-shot-image-retrieval","task_name":"Zero-shot Image Retrieval"},{"task_slug":"zero-shot-text-to-image-retrieval","task_name":"Zero-shot Text-to-Image Retrieval"},{"task_slug":null,"task_name":"zero-shot-classification"}],"methods":[{"method_slug":"clip","method_name":"CLIP"},{"method_slug":"set","method_name":"SET"},{"method_slug":"sycoca","method_name":"SyCoCa"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/zero-shot-cross-modal-retrieval-on-coco-2014","task":"Zero-Shot Cross-Modal Retrieval","dataset":"COCO 2014","model":"M2-Encoder","rank_in_archive_order":2,"of":18,"metrics":{"Image-to-text R@1":"72.8","Image-to-text R@10":"96.3","Image-to-text R@5":"92.3","Text-to-image R@1":"56.5","Text-to-image R@10":"88.8","Text-to-image R@5":"81.6"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-cross-modal-retrieval-on-flickr30k","task":"Zero-Shot Cross-Modal Retrieval","dataset":"Flickr30k","model":"M2-Encoder","rank_in_archive_order":7,"of":22,"metrics":{"Image-to-text R@1":"91.2","Image-to-text R@10":"99.6","Image-to-text R@5":"99.2","Text-to-image R@1":"92.2","Text-to-image R@10":"99.7","Text-to-image R@5":"99.5"},"uses_additional_data":true},{"leaderboard":"/sota/zero-shot-learning-on-imagenet-cn","task":"Zero-Shot Learning","dataset":"ImageNet_CN","model":"$M^2$-Encoder","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy":"80.7"},"uses_additional_data":true},{"leaderboard":"/sota/zero-shot-transfer-image-classification-on-1","task":"Zero-Shot Transfer Image Classification","dataset":"ImageNet","model":"M2-Encoder","rank_in_archive_order":1,"of":23,"metrics":{"Accuracy (Private)":"88.5","Param":"10B"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2401.15896","atlas_url":"https://app.syntology.ai/?focus=2401.15896","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}