Methods › Natural Language Processing › Word Embeddings › UNITER

UNiversal Image-TExt Representation Learning

UNITER

23 papers tagged archive 2025-07-28

Introduced by Yen-Chun Chen et al. in UNITER: UNiversal Image-TExt Representation Learning

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

UNITER or UNiversal Image-TExt Representation model is a large-scale pre-trained model for joint multimodal embedding. It is pre-trained using four image-text datasets COCO, Visual Genome, Conceptual Captions, and SBU Captions. It can power heterogeneous downstream V+L tasks with joint multimodal embeddings. UNITER takes the visual regions of the image and textual tokens of the sentence as inputs. A faster R-CNN is used in Image Embedder to extract the visual features of each region and a Text Embedder is used to tokenize the input sentence into WordPieces.

It proposes WRA via the Optimal Transport to provide more fine-grained alignment between word tokens and image regions that is effective in calculating the minimum cost of transporting the contextualized image embeddings to word embeddings and vice versa.

Four pretraining tasks were designed for this model. They are Masked Language Modeling (MLM), Masked Region Modeling (MRM, with three variants), Image-Text Matching (ITM), and Word-Region Alignment (WRA). This model is different from the previous models because it uses conditional masking on pre-training tasks.

PaperSource

Papers archive 2025-07-28

23 shown of 23, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 39 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Question Answering6
Referring Expression6
Referring Expression Comprehension6
Retrieval6
Visual Question Answering6
Image-text Retrieval5
Text Retrieval5
Visual Question Answering (VQA)5
Language Modelling4
Image Captioning3
Image-text matching3
Language Modeling3
Representation Learning3
Text Matching3
Visual Commonsense Reasoning3
Visual Entailment3
Contrastive Learning2
Data Augmentation2
GPU2
Masked Language Modeling2

Usage over time archive 2025-07-28

Papers per year tagged with UNITER: 2019 to 2024, peak 8 8 0 2019: 2 papers 2019 2020: 4 papers 2020 2021: 6 papers 2021 2022: 8 papers 2022 2023: 2 papers 2023 2024: 1 paper 2024
Papers per year the archive tags with this method, by the paper's archive date (23 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Word Embeddings

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections