{"url":"/method/cmcl","slug":"cmcl","name":"CMCL","full_name":"Crossmodal Contrastive Learning","full_name_withheld":false,"description_markdown":"**CMCL**, or **Crossmodal Contrastive Learning**, is a method for unifying visual and textual representations into the same semantic space based on a large-scale corpus of image collections, text corpus and image-text pairs. The CMCL aligns the visual representations and textual representations, and unifies them into the same semantic space based on image-text pairs. As shown in the Figure, to facilitate different levels of semantic alignment between vision and language, a series of text rewriting techniques are utilized to improve the diversity of cross-modal information. Specifically, for an image-text pair, various positive examples and hard negative examples can be obtained by rewriting the original caption at different levels. Moreover, to incorporate more background information from the single-modal data, text and image retrieval are also applied to augment each image-text pair with various related texts and images. The positive pairs, negative pairs, related images and texts are learned jointly by CMCL. In this way, the model can effectively unify different levels of visual and textual representations into the same semantic space, and incorporate more single-modal knowledge to enhance each other.","description_state":"present","introduced_year":null,"introduced_by":{"title":"UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning","paper":"/paper/unimo-towards-unified-modal-understanding-and","first_author":"Wei Li","n_authors":8,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/unimo-towards-unified-modal-understanding-and"},"source":{"url":"https://arxiv.org/abs/2012.15409v4","title":"UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Self-Supervised Learning","url":"/methods/category/self-supervised-learning","pwc_aliases":[]}],"n_papers_tagged":12,"archive_num_papers":12,"papers_newest_first":[{"paper":"/paper/consistency-aware-fake-videos-detection-on","title":"Consistency-aware Fake Videos Detection on Short Video Platforms","date":"2025-04-30","arxiv_id":"2504.21495","n_code_links":1,"syntology":null},{"paper":"/paper/continual-multimodal-contrastive-learning","title":"Continual Multimodal Contrastive Learning","date":"2025-03-19","arxiv_id":"2503.14963","n_code_links":0,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":1}},{"paper":"/paper/call-for-papers-the-babylm-challenge-sample","title":"Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus","date":"2023-01-27","arxiv_id":"2301.11796","n_code_links":1,"syntology":{"ran":0,"of":1,"unverified":1,"pointer_only":1}},{"paper":null,"title":"Competence-based Multimodal Curriculum Learning for Medical Report Generation","date":"2022-06-24","arxiv_id":"2206.14579","n_code_links":0,"syntology":null},{"paper":null,"title":"Team ÚFAL at CMCL 2022 Shared Task: Figuring out the correct recipe for predicting Eye-Tracking features using Pretrained Language Models","date":"2022-04-11","arxiv_id":"2204.04998","n_code_links":0,"syntology":null},{"paper":null,"title":"Zero Shot Crosslingual Eye-Tracking Data Prediction using Multilingual Transformer Models","date":"2022-03-30","arxiv_id":"2203.16474","n_code_links":0,"syntology":null},{"paper":null,"title":"WuDaoMM: A large-scale Multi-Modal Dataset for Pre-training models","date":"2022-03-22","arxiv_id":"2203.11480","n_code_links":0,"syntology":null},{"paper":"/paper/unimo-2-end-to-end-unified-vision-language","title":"UNIMO-2: End-to-End Unified Vision-Language Grounded Learning","date":"2022-03-17","arxiv_id":"2203.09067","n_code_links":1,"syntology":null},{"paper":null,"title":"A Multimodal Sentiment Dataset for Video Recommendation","date":"2021-09-17","arxiv_id":"2109.08333","n_code_links":0,"syntology":null},{"paper":null,"title":"LAST at CMCL 2021 Shared Task: Predicting Gaze Data During Reading with a Gradient Boosting Decision Tree Approach","date":"2021-04-27","arxiv_id":"2104.13043","n_code_links":0,"syntology":null},{"paper":"/paper/torontocl-at-cmcl-2021-shared-task-roberta","title":"TorontoCL at CMCL 2021 Shared Task: RoBERTa with Multi-Stage Fine-Tuning for Eye-Tracking Prediction","date":"2021-04-15","arxiv_id":"2104.07244","n_code_links":1,"syntology":null},{"paper":"/paper/unimo-towards-unified-modal-understanding-and","title":"UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning","date":"2020-12-31","arxiv_id":"2012.15409","n_code_links":3,"syntology":null}],"papers_shown":12,"tasks":[{"task":"/task/image-captioning","name":"Image Captioning","papers":3},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":2},{"task":"/task/all","name":"All","papers":1},{"task":"/task/continual-learning","name":"Continual Learning","papers":1},{"task":"/task/cross-modal-retrieval","name":"Cross-Modal Retrieval","papers":1},{"task":"/task/image-generation","name":"Image Generation","papers":1},{"task":"/task/language-acquisition","name":"Language Acquisition","papers":1},{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":"/task/large-language-model","name":"Large Language Model","papers":1},{"task":"/task/medical-report-generation","name":"Medical Report Generation","papers":1},{"task":"/task/multimodal-large-language-model","name":"Multimodal Large Language Model","papers":1},{"task":"/task/multimodal-sentiment-analysis","name":"Multimodal Sentiment Analysis","papers":1},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":1},{"task":"/task/pretrained-multilingual-language-models","name":"Pretrained Multilingual Language Models","papers":1},{"task":"/task/pseudo-label","name":"Pseudo Label","papers":1},{"task":"/task/question-answering","name":"Question Answering","papers":1},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":1},{"task":"/task/text-to-image-generation-1","name":"Text to Image Generation","papers":1},{"task":"/task/text-to-image-generation","name":"Text-to-Image Generation","papers":1}],"tasks_shown":20,"n_tasks":23,"usage_by_year":[{"year":"2020","papers":1},{"year":"2021","papers":3},{"year":"2022","papers":5},{"year":"2023","papers":1},{"year":"2025","papers":2}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/cmcl"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}