Methods › Computer Vision › Vision and Language Pre-Trained Models › UNIMO

UNIMO

4 papers tagged archive 2025-07-28

Introduced by Wei Li et al. in UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

UNIMO is a multi-modal pre-training architecture that can effectively adapt to both single modal and multimodal understanding and generation tasks. UNIMO learns visual representations and textual representations simultaneously, and unifies them into the same semantic space via cross-modal contrastive learning (CMCL) based on a large-scale corpus of image collections, text corpus and image-text pairs. The CMCL aligns the visual representation and textual representation, and unifies them into the same semantic space based on image-text pairs.

PaperSource

Papers archive 2025-07-28

4 shown of 4, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

11 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Image Captioning2
Contrastive Learning1
Cross-Modal Retrieval1
Image Generation1
Multimodal Sentiment Analysis1
Question Answering1
Sentiment Analysis1
Text to Image Generation1
Text-to-Image Generation1
Video Understanding1
Visual Question Answering (VQA)1

Usage over time archive 2025-07-28

Papers per year tagged with UNIMO: 2020 to 2022, peak 2 2 0 2020: 1 paper 2020 2021: 1 paper 2021 2022: 2 papers 2022
Papers per year the archive tags with this method, by the paper's archive date (4 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Vision and Language Pre-Trained ModelsMulti-Modal Methods

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections