Methods › General › Self-Supervised Learning › CMCL
Crossmodal Contrastive Learning
CMCL
Introduced by Wei Li et al. in UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
CMCL, or Crossmodal Contrastive Learning, is a method for unifying visual and textual representations into the same semantic space based on a large-scale corpus of image collections, text corpus and image-text pairs. The CMCL aligns the visual representations and textual representations, and unifies them into the same semantic space based on image-text pairs. As shown in the Figure, to facilitate different levels of semantic alignment between vision and language, a series of text rewriting techniques are utilized to improve the diversity of cross-modal information. Specifically, for an image-text pair, various positive examples and hard negative examples can be obtained by rewriting the original caption at different levels. Moreover, to incorporate more background information from the single-modal data, text and image retrieval are also applied to augment each image-text pair with various related texts and images. The positive pairs, negative pairs, related images and texts are learned jointly by CMCL. In this way, the model can effectively unify different levels of visual and textual representations into the same semantic space, and incorporate more single-modal knowledge to enhance each other.
Papers archive 2025-07-28
12 shown of 12, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Consistency-aware Fake Videos Detection on Short Video Platforms 30 Apr 2025 · 1 repository · arXiv:2504.21495
-
Continual Multimodal Contrastive Learning 19 Mar 2025 · 0 repositories · arXiv:2503.14963Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)
-
Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus 27 Jan 2023 · 1 repository · arXiv:2301.11796Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)
-
Competence-based Multimodal Curriculum Learning for Medical Report Generation 24 Jun 2022 · 0 repositories · arXiv:2206.14579
-
Team ÚFAL at CMCL 2022 Shared Task: Figuring out the correct recipe for predicting Eye-Tracking features using Pretrained Language Models 11 Apr 2022 · 0 repositories · arXiv:2204.04998
-
Zero Shot Crosslingual Eye-Tracking Data Prediction using Multilingual Transformer Models 30 Mar 2022 · 0 repositories · arXiv:2203.16474
-
WuDaoMM: A large-scale Multi-Modal Dataset for Pre-training models 22 Mar 2022 · 0 repositories · arXiv:2203.11480
-
UNIMO-2: End-to-End Unified Vision-Language Grounded Learning 17 Mar 2022 · 1 repository · arXiv:2203.09067
-
A Multimodal Sentiment Dataset for Video Recommendation 17 Sep 2021 · 0 repositories · arXiv:2109.08333
-
LAST at CMCL 2021 Shared Task: Predicting Gaze Data During Reading with a Gradient Boosting Decision Tree Approach 27 Apr 2021 · 0 repositories · arXiv:2104.13043
-
TorontoCL at CMCL 2021 Shared Task: RoBERTa with Multi-Stage Fine-Tuning for Eye-Tracking Prediction 15 Apr 2021 · 1 repository · arXiv:2104.07244
-
UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning 31 Dec 2020 · 3 repositories · arXiv:2012.15409
Tasks archive 2025-07-28
20 shown of 23 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections