Methods › Natural Language Processing › Word Embeddings › UNITER
UNiversal Image-TExt Representation Learning
UNITER
Introduced by Yen-Chun Chen et al. in UNITER: UNiversal Image-TExt Representation Learning
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
UNITER or UNiversal Image-TExt Representation model is a large-scale pre-trained model for joint multimodal embedding. It is pre-trained using four image-text datasets COCO, Visual Genome, Conceptual Captions, and SBU Captions. It can power heterogeneous downstream V+L tasks with joint multimodal embeddings. UNITER takes the visual regions of the image and textual tokens of the sentence as inputs. A faster R-CNN is used in Image Embedder to extract the visual features of each region and a Text Embedder is used to tokenize the input sentence into WordPieces.
It proposes WRA via the Optimal Transport to provide more fine-grained alignment between word tokens and image regions that is effective in calculating the minimum cost of transporting the contextualized image embeddings to word embeddings and vice versa.
Four pretraining tasks were designed for this model. They are Masked Language Modeling (MLM), Masked Region Modeling (MRM, with three variants), Image-Text Matching (ITM), and Word-Region Alignment (WRA). This model is different from the previous models because it uses conditional masking on pre-training tasks.
Papers archive 2025-07-28
23 shown of 23, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking 29 Jan 2024 · 1 repository · arXiv:2401.16575
-
Switching Head-Tail Funnel UNITER for Dual Referring Expression Comprehension with Fetch-and-Carry Tasks 14 Jul 2023 · 0 repositories · arXiv:2307.07166
-
Switch-BERT: Learning to Model Multimodal Interactions by Switching Attention and Input 25 Jun 2023 · 0 repositories · arXiv:2306.14182
-
Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment 20 Dec 2022 · 1 repository · arXiv:2212.10549
-
Probing Cross-modal Semantics Alignment Capability from the Textual Perspective 18 Oct 2022 · 0 repositories · arXiv:2210.09550
-
VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations 1 Jul 2022 · 1 repository · arXiv:2207.00221
-
Entity-Graph Enhanced Cross-Modal Pretraining for Instance-level Product Retrieval 17 Jun 2022 · 0 repositories · arXiv:2206.08842
-
UPB at SemEval-2022 Task 5: Enhancing UNITER with Image Sentiment and Graph Convolutional Networks for Multimedia Automatic Misogyny Identification 29 May 2022 · 1 repository · arXiv:2205.14769
-
HiVLP: Hierarchical Vision-Language Pre-Training for Fast Image-Text Retrieval 24 May 2022 · 0 repositories · arXiv:2205.12105
-
A Survivor in the Era of Large-Scale Pretraining: An Empirical Study of One-Stage Referring Expression Comprehension 17 Apr 2022 · 1 repository · arXiv:2204.07913
-
Bilaterally Slimmable Transformer for Elastic and Efficient Visual Question Answering 24 Mar 2022 · 1 repository · arXiv:2203.12814
-
Hateful Memes Challenge: An Enhanced Multimodal Framework 20 Dec 2021 · 1 repository · arXiv:2112.11244
-
Dense Contrastive Visual-Linguistic Pretraining 24 Sep 2021 · 0 repositories · arXiv:2109.11778
-
Target-dependent UNITER: A Transformer-Based Multimodal Language Comprehension Model for Domestic Service Robots 2 Jul 2021 · 0 repositories · arXiv:2107.00811
-
e-ViL: A Dataset and Benchmark for Natural Language Explanations in Vision-Language Tasks 8 May 2021 · 2 repositories · arXiv:2105.03761Syntology ran 7 of 15 samples · 8 unverified · 15 pointer-only (licence)
-
Playing Lottery Tickets with Vision and Language 23 Apr 2021 · 0 repositories · arXiv:2104.11832
-
WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training 11 Mar 2021 · 2 repositories · arXiv:2103.06561
-
A Closer Look at the Robustness of Vision-and-Language Pre-trained Models 15 Dec 2020 · 0 repositories · arXiv:2012.08673
-
X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers 23 Sep 2020 · 2 repositories · arXiv:2009.11278Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)
-
What Does BERT with Vision Look At? 1 Jul 2020 · 0 repositories
-
Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models 15 May 2020 · 0 repositories · arXiv:2005.07310
-
UNITER: Learning UNiversal Image-TExt Representations 25 Sep 2019 · 0 repositories
-
UNITER: UNiversal Image-TExt Representation Learning 25 Sep 2019 · 7 repositories · arXiv:1909.11740Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 39 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections