Methods › General › Self-Supervised Learning › CMCL

Crossmodal Contrastive Learning

CMCL

12 papers tagged archive 2025-07-28

Introduced by Wei Li et al. in UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

CMCL, or Crossmodal Contrastive Learning, is a method for unifying visual and textual representations into the same semantic space based on a large-scale corpus of image collections, text corpus and image-text pairs. The CMCL aligns the visual representations and textual representations, and unifies them into the same semantic space based on image-text pairs. As shown in the Figure, to facilitate different levels of semantic alignment between vision and language, a series of text rewriting techniques are utilized to improve the diversity of cross-modal information. Specifically, for an image-text pair, various positive examples and hard negative examples can be obtained by rewriting the original caption at different levels. Moreover, to incorporate more background information from the single-modal data, text and image retrieval are also applied to augment each image-text pair with various related texts and images. The positive pairs, negative pairs, related images and texts are learned jointly by CMCL. In this way, the model can effectively unify different levels of visual and textual representations into the same semantic space, and incorporate more single-modal knowledge to enhance each other.

PaperSource

Papers archive 2025-07-28

12 shown of 12, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 23 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Image Captioning3
Contrastive Learning2
All1
Continual Learning1
Cross-Modal Retrieval1
Image Generation1
Language Acquisition1
Language Modeling1
Language Modelling1
Large Language Model1
Medical Report Generation1
Multimodal Large Language Model1
Multimodal Sentiment Analysis1
Natural Language Understanding1
Pretrained Multilingual Language Models1
Pseudo Label1
Question Answering1
Sentiment Analysis1
Text to Image Generation1
Text-to-Image Generation1

Usage over time archive 2025-07-28

Papers per year tagged with CMCL: 2020 to 2025, peak 5 5 0 2020: 1 paper 2020 2021: 3 papers 2021 2022: 5 papers 2022 2023: 1 paper 2023 2024: 0 papers 2024 2025: 2 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (12 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Self-Supervised Learning

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections