{"url":"/method/uniter","slug":"uniter","name":"UNITER","full_name":"UNiversal Image-TExt Representation Learning","full_name_withheld":false,"description_markdown":"UNITER or UNiversal Image-TExt Representation model is a large-scale pre-trained model for joint multimodal embedding. It is pre-trained using four image-text datasets COCO, Visual Genome, Conceptual Captions, and SBU Captions. It can power heterogeneous downstream V+L tasks with joint multimodal embeddings. \r\nUNITER takes the visual regions of the image and textual tokens of the sentence as inputs. A faster R-CNN is used in Image Embedder to extract the visual features of each region and a Text Embedder is used to tokenize the input sentence into WordPieces.  \r\n\r\nIt proposes WRA via the Optimal Transport to provide more fine-grained alignment between word tokens and image regions that is effective in calculating the minimum cost of transporting the contextualized image embeddings to word embeddings and vice versa. \r\n\r\nFour pretraining tasks were designed for this model. They are Masked Language Modeling (MLM), Masked Region Modeling (MRM, with three variants), Image-Text Matching (ITM), and Word-Region Alignment (WRA). This model is different from the previous models because it uses conditional masking on pre-training tasks.","description_state":"present","introduced_year":null,"introduced_by":{"title":"UNITER: UNiversal Image-TExt Representation Learning","paper":"/paper/uniter-learning-universal-image-text-1","first_author":"Yen-Chun Chen","n_authors":8,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/uniter-learning-universal-image-text-1"},"source":{"url":"https://arxiv.org/abs/1909.11740v3","title":"UNITER: UNiversal Image-TExt Representation Learning","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Word Embeddings","url":"/methods/category/word-embeddings","pwc_aliases":[]}],"n_papers_tagged":23,"archive_num_papers":23,"papers_newest_first":[{"paper":"/paper/beyond-image-text-matching-verb-understanding","title":"Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking","date":"2024-01-29","arxiv_id":"2401.16575","n_code_links":1,"syntology":null},{"paper":null,"title":"Switching Head-Tail Funnel UNITER for Dual Referring Expression Comprehension with Fetch-and-Carry Tasks","date":"2023-07-14","arxiv_id":"2307.07166","n_code_links":0,"syntology":null},{"paper":null,"title":"Switch-BERT: Learning to Model Multimodal Interactions by Switching Attention and Input","date":"2023-06-25","arxiv_id":"2306.14182","n_code_links":0,"syntology":null},{"paper":"/paper/cross-modal-attention-congruence","title":"Cross-modal Attention Congruence Regularization for Vision-Language Relation Alignment","date":"2022-12-20","arxiv_id":"2212.10549","n_code_links":1,"syntology":null},{"paper":null,"title":"Probing Cross-modal Semantics Alignment Capability from the Textual Perspective","date":"2022-10-18","arxiv_id":"2210.09550","n_code_links":0,"syntology":null},{"paper":"/paper/vl-checklist-evaluating-pre-trained-vision","title":"VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations","date":"2022-07-01","arxiv_id":"2207.00221","n_code_links":1,"syntology":null},{"paper":null,"title":"Entity-Graph Enhanced Cross-Modal Pretraining for Instance-level Product Retrieval","date":"2022-06-17","arxiv_id":"2206.08842","n_code_links":0,"syntology":null},{"paper":"/paper/upb-at-semeval-2022-task-5-enhancing-uniter","title":"UPB at SemEval-2022 Task 5: Enhancing UNITER with Image Sentiment and Graph Convolutional Networks for Multimedia Automatic Misogyny Identification","date":"2022-05-29","arxiv_id":"2205.14769","n_code_links":1,"syntology":null},{"paper":null,"title":"HiVLP: Hierarchical Vision-Language Pre-Training for Fast Image-Text Retrieval","date":"2022-05-24","arxiv_id":"2205.12105","n_code_links":0,"syntology":null},{"paper":"/paper/what-goes-beyond-multi-modal-fusion-in-one","title":"A Survivor in the Era of Large-Scale Pretraining: An Empirical Study of One-Stage Referring Expression Comprehension","date":"2022-04-17","arxiv_id":"2204.07913","n_code_links":1,"syntology":null},{"paper":"/paper/towards-efficient-and-elastic-visual-question","title":"Bilaterally Slimmable Transformer for Elastic and Efficient Visual Question Answering","date":"2022-03-24","arxiv_id":"2203.12814","n_code_links":1,"syntology":null},{"paper":"/paper/hateful-memes-challenge-an-enhanced","title":"Hateful Memes Challenge: An Enhanced Multimodal Framework","date":"2021-12-20","arxiv_id":"2112.11244","n_code_links":1,"syntology":null},{"paper":null,"title":"Dense Contrastive Visual-Linguistic Pretraining","date":"2021-09-24","arxiv_id":"2109.11778","n_code_links":0,"syntology":null},{"paper":null,"title":"Target-dependent UNITER: A Transformer-Based Multimodal Language Comprehension Model for Domestic Service Robots","date":"2021-07-02","arxiv_id":"2107.00811","n_code_links":0,"syntology":null},{"paper":"/paper/e-vil-a-dataset-and-benchmark-for-natural","title":"e-ViL: A Dataset and Benchmark for Natural Language Explanations in Vision-Language Tasks","date":"2021-05-08","arxiv_id":"2105.03761","n_code_links":2,"syntology":{"ran":7,"of":15,"unverified":8,"pointer_only":15}},{"paper":null,"title":"Playing Lottery Tickets with Vision and Language","date":"2021-04-23","arxiv_id":"2104.11832","n_code_links":0,"syntology":null},{"paper":"/paper/wenlan-bridging-vision-and-language-by-large","title":"WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training","date":"2021-03-11","arxiv_id":"2103.06561","n_code_links":2,"syntology":null},{"paper":null,"title":"A Closer Look at the Robustness of Vision-and-Language Pre-trained Models","date":"2020-12-15","arxiv_id":"2012.08673","n_code_links":0,"syntology":null},{"paper":"/paper/x-lxmert-paint-caption-and-answer-questions","title":"X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers","date":"2020-09-23","arxiv_id":"2009.11278","n_code_links":2,"syntology":{"ran":1,"of":2,"unverified":1,"pointer_only":2}},{"paper":null,"title":"What Does BERT with Vision Look At?","date":"2020-07-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models","date":"2020-05-15","arxiv_id":"2005.07310","n_code_links":0,"syntology":null},{"paper":null,"title":"UNITER: Learning UNiversal Image-TExt Representations","date":"2019-09-25","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/uniter-learning-universal-image-text-1","title":"UNITER: UNiversal Image-TExt Representation Learning","date":"2019-09-25","arxiv_id":"1909.11740","n_code_links":7,"syntology":{"ran":3,"of":3,"unverified":0,"pointer_only":2}}],"papers_shown":23,"tasks":[{"task":"/task/question-answering","name":"Question Answering","papers":6},{"task":"/task/referring-expression","name":"Referring Expression","papers":6},{"task":"/task/referring-expression-comprehension","name":"Referring Expression Comprehension","papers":6},{"task":"/task/retrieval","name":"Retrieval","papers":6},{"task":"/task/visual-question-answering-1","name":"Visual Question Answering","papers":6},{"task":"/task/image-text-retrieval","name":"Image-text Retrieval","papers":5},{"task":"/task/text-retrieval","name":"Text Retrieval","papers":5},{"task":"/task/visual-question-answering","name":"Visual Question Answering (VQA)","papers":5},{"task":"/task/language-modelling","name":"Language Modelling","papers":4},{"task":"/task/image-captioning","name":"Image Captioning","papers":3},{"task":"/task/image-text-matching","name":"Image-text matching","papers":3},{"task":"/task/language-modeling","name":"Language Modeling","papers":3},{"task":"/task/representation-learning","name":"Representation Learning","papers":3},{"task":"/task/text-matching","name":"Text Matching","papers":3},{"task":"/task/visual-commonsense-reasoning","name":"Visual Commonsense Reasoning","papers":3},{"task":"/task/visual-entailment","name":"Visual Entailment","papers":3},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":2},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":2},{"task":null,"name":"GPU","papers":2},{"task":"/task/masked-language-modeling","name":"Masked Language Modeling","papers":2}],"tasks_shown":20,"n_tasks":39,"usage_by_year":[{"year":"2019","papers":2},{"year":"2020","papers":4},{"year":"2021","papers":6},{"year":"2022","papers":8},{"year":"2023","papers":2},{"year":"2024","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/uniter"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}