{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ernie-vil-2-0-multi-view-contrastive-learning","title":"ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training","arxiv_id":"2209.15270","date":"2022-09-30","proceeding":null,"authors":["Bin Shan","Weichong Yin","Yu Sun","Hao Tian","Hua Wu","Haifeng Wang"],"abstract":"Recent Vision-Language Pre-trained (VLP) models based on dual encoder have attracted extensive attention from academia and industry due to their superior performance on various cross-modal tasks and high computational efficiency. They attempt to learn cross-modal representation using contrastive learning on image-text pairs, however, the built inter-modal correlations only rely on a single view for each modality. Actually, an image or a text contains various potential views, just as humans could capture a real-world scene via diverse descriptions or photos. In this paper, we propose ERNIE-ViL 2.0, a Multi-View Contrastive learning framework to build intra-modal and inter-modal correlations between diverse views simultaneously, aiming at learning a more robust cross-modal representation. Specifically, we construct multiple views within each modality to learn the intra-modal correlation for enhancing the single-modal representation. Besides the inherent visual/textual views, we construct sequences of object tags as a special textual view to narrow the cross-modal semantic gap on noisy image-text pairs. Pre-trained with 29M publicly available datasets, ERNIE-ViL 2.0 achieves competitive results on English cross-modal retrieval. Additionally, to generalize our method to Chinese cross-modal tasks, we train ERNIE-ViL 2.0 through scaling up the pre-training datasets to 1.5B Chinese image-text pairs, resulting in significant improvements compared to previous SOTA results on Chinese cross-modal retrieval. We release our pre-trained models in https://github.com/PaddlePaddle/ERNIE.","url_abs":"https://arxiv.org/abs/2209.15270v1","url_pdf":"https://arxiv.org/pdf/2209.15270v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ernie-vil-2-0-multi-view-contrastive-learning","repo_url":"https://github.com/PaddlePaddle/ERNIE","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"paddle","reach":null}],"tasks":[{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"},{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"},{"task_slug":"cross-modal-retrieval","task_name":"Cross-Modal Retrieval"},{"task_slug":"image-retrieval","task_name":"Image Retrieval"},{"task_slug":"image-to-text-retrieval","task_name":"Image-to-Text Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"zero-shot-cross-modal-retrieval","task_name":"Zero-Shot Cross-Modal Retrieval"},{"task_slug":"zero-shot-image-retrieval","task_name":"Zero-shot Image Retrieval"},{"task_slug":"zero-shot-text-to-image-retrieval","task_name":"Zero-shot Text-to-Image Retrieval"}],"methods":[{"method_slug":"contrastive-learning","method_name":"Contrastive Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/cross-modal-retrieval-on-coco-2014","task":"Cross-Modal Retrieval","dataset":"COCO 2014","model":"ERNIE-ViL 2.0","rank_in_archive_order":17,"of":36,"metrics":{"Image-to-text R@1":"77.4","Image-to-text R@10":"97.1","Image-to-text R@5":"93.6","Text-to-image R@1":"59.5","Text-to-image R@10":"90.1","Text-to-image R@5":"83.4"},"uses_additional_data":true},{"leaderboard":"/sota/cross-modal-retrieval-on-flickr30k","task":"Cross-Modal Retrieval","dataset":"Flickr30k","model":"ERNIE-ViL 2.0","rank_in_archive_order":5,"of":27,"metrics":{"Image-to-text R@1":"97.2","Image-to-text R@10":"100.0","Image-to-text R@5":"100.0","Text-to-image R@1":"93.3","Text-to-image R@10":"99.8","Text-to-image R@5":"99.4"},"uses_additional_data":true},{"leaderboard":"/sota/image-retrieval-on-aic-icc","task":"Image Retrieval","dataset":"AIC-ICC","model":"ERNIE-ViL2.0","rank_in_archive_order":1,"of":2,"metrics":{"Recall@1":"19.0","Recall@10":"43.5","Recall@5":"35.3"},"uses_additional_data":false},{"leaderboard":"/sota/image-to-text-retrieval-on-aic-icc","task":"Image-to-Text Retrieval","dataset":"AIC-ICC","model":"ERNIE-ViL2.0","rank_in_archive_order":1,"of":2,"metrics":{"Recall@1":"33.7","Recall@10":"60.0","Recall@5":"52.1"},"uses_additional_data":false},{"leaderboard":"/sota/image-to-text-retrieval-on-flickr30k","task":"Image-to-Text Retrieval","dataset":"Flickr30k","model":"ERNIE-ViL 2.0","rank_in_archive_order":6,"of":11,"metrics":{"Recall@1":"96.1","Recall@10":"100.0","Recall@5":"99.9"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-cross-modal-retrieval-on-coco-2014","task":"Zero-Shot Cross-Modal Retrieval","dataset":"COCO 2014","model":"ERNIE-ViL 2.0","rank_in_archive_order":13,"of":18,"metrics":{"Image-to-text R@1":"63.1","Image-to-text R@10":"91.4","Image-to-text R@5":"85.7","Text-to-image R@1":"46.0","Text-to-image R@10":"80.4","Text-to-image R@5":"71.4"},"uses_additional_data":false},{"leaderboard":"/sota/zero-shot-cross-modal-retrieval-on-flickr30k","task":"Zero-Shot Cross-Modal Retrieval","dataset":"Flickr30k","model":"ERNIE-ViL 2.0","rank_in_archive_order":8,"of":22,"metrics":{"Image-to-text R@1":"91.2","Image-to-text R@10":"99.8","Image-to-text R@5":"99.1","Text-to-image R@1":"77.4","Text-to-image R@10":"96.4","Text-to-image R@5":"93.8"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2209.15270","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}