{"url":"/method/lxmert","slug":"lxmert","name":"LXMERT","full_name":"Learning Cross-Modality Encoder Representations from Transformers","full_name_withheld":false,"description_markdown":"LXMERT is a model for learning vision-and-language cross-modality representations. It consists of a Transformer model that consists three encoders: object relationship encoder, a language encoder, and a cross-modality encoder. The model takes two inputs: image with its related sentence. The images are represented as a sequence of objects, whereas each sentence is represented as sequence of words. By combining the self-attention and cross-attention layers the model is able to generated language representation, image representations, and cross-modality representations from the input. The model is pre-trained with image-sentence pairs via five pre-training tasks: masked language modeling, masked object prediction, cross-modality matching, and image questions answering. These tasks help the model to learn both intra-modality and cross-modality relationships.","description_state":"present","introduced_year":null,"introduced_by":{"title":"LXMERT: Learning Cross-Modality Encoder Representations from Transformers","paper":"/paper/lxmert-learning-cross-modality-encoder","first_author":"Hao Tan","n_authors":2,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/lxmert-learning-cross-modality-encoder"},"source":{"url":"https://arxiv.org/abs/1908.07490v3","title":"LXMERT: Learning Cross-Modality Encoder Representations from Transformers","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision and Language Pre-Trained Models","url":"/methods/category/vision-and-language-pre-trained-models","pwc_aliases":[]}],"n_papers_tagged":40,"archive_num_papers":40,"papers_newest_first":[{"paper":null,"title":"Optimizing Visual Question Answering Models for Driving: Bridging the Gap Between Human and Machine Attention Patterns","date":"2024-06-13","arxiv_id":"2406.09203","n_code_links":0,"syntology":null},{"paper":"/paper/beyond-image-text-matching-verb-understanding","title":"Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking","date":"2024-01-29","arxiv_id":"2401.16575","n_code_links":1,"syntology":null},{"paper":"/paper/lxmert-model-compression-for-visual-question","title":"LXMERT Model Compression for Visual Question Answering","date":"2023-10-23","arxiv_id":"2310.15325","n_code_links":2,"syntology":null},{"paper":null,"title":"Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models","date":"2023-08-18","arxiv_id":"2308.09778","n_code_links":0,"syntology":null},{"paper":"/paper/towards-a-performance-analysis-on-pre-trained","title":"Towards a performance analysis on pre-trained Visual Question Answering models for autonomous driving","date":"2023-07-18","arxiv_id":"2307.09329","n_code_links":1,"syntology":null},{"paper":null,"title":"An Empirical Study on the Language Modal in Visual Question Answering","date":"2023-05-17","arxiv_id":"2305.10143","n_code_links":0,"syntology":null},{"paper":null,"title":"Probing the Role of Positional Information in Vision-Language Models","date":"2023-05-17","arxiv_id":"2305.10046","n_code_links":0,"syntology":null},{"paper":null,"title":"Controlling for Stereotypes in Multimodal Language Model Evaluation","date":"2023-02-03","arxiv_id":"2302.01582","n_code_links":0,"syntology":null},{"paper":"/paper/mm-shap-a-performance-agnostic-metric-for","title":"MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks","date":"2022-12-15","arxiv_id":"2212.08158","n_code_links":1,"syntology":{"ran":1,"of":10,"unverified":9,"pointer_only":0}},{"paper":null,"title":"Enhancing Self-Consistency and Performance of Pre-Trained Language Models through Natural Language Inference","date":"2022-11-21","arxiv_id":"2211.11875","n_code_links":0,"syntology":null},{"paper":"/paper/compressing-and-debiasing-vision-language-pre","title":"Compressing And Debiasing Vision-Language Pre-Trained Models for Visual Question Answering","date":"2022-10-26","arxiv_id":"2210.14558","n_code_links":1,"syntology":null},{"paper":null,"title":"Probing Cross-modal Semantics Alignment Capability from the Textual Perspective","date":"2022-10-18","arxiv_id":"2210.09550","n_code_links":0,"syntology":null},{"paper":"/paper/generative-bias-for-visual-question-answering","title":"Generative Bias for Robust Visual Question Answering","date":"2022-08-01","arxiv_id":"2208.00690","n_code_links":1,"syntology":{"ran":0,"of":2,"unverified":2,"pointer_only":0}},{"paper":"/paper/vl-checklist-evaluating-pre-trained-vision","title":"VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations","date":"2022-07-01","arxiv_id":"2207.00221","n_code_links":1,"syntology":null},{"paper":"/paper/visual-spatial-reasoning","title":"Visual Spatial Reasoning","date":"2022-04-30","arxiv_id":"2205.00363","n_code_links":4,"syntology":{"ran":5,"of":10,"unverified":5,"pointer_only":0}},{"paper":null,"title":"Visio-Linguistic Brain Encoding","date":"2022-04-18","arxiv_id":"2204.08261","n_code_links":0,"syntology":null},{"paper":"/paper/swapmix-diagnosing-and-regularizing-the-over","title":"SwapMix: Diagnosing and Regularizing the Over-Reliance on Visual Context in Visual Question Answering","date":"2022-04-05","arxiv_id":"2204.02285","n_code_links":1,"syntology":null},{"paper":"/paper/exploring-multi-modal-representations-for","title":"Exploring Multi-Modal Representations for Ambiguity Detection & Coreference Resolution in the SIMMC 2.0 Challenge","date":"2022-02-25","arxiv_id":"2202.12645","n_code_links":2,"syntology":null},{"paper":null,"title":"Probing the Role of Positional Information in Vision-Language Models","date":"2022-01-16","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Multimodal Learning: Are Captions All You Need?","date":"2021-11-16","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Region under Discussion for visual dialog","date":"2021-11-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Seeing things or seeing scenes: Investigating the capabilities of V&L models to align scene descriptions to images","date":"2021-10-16","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Dense Contrastive Visual-Linguistic Pretraining","date":"2021-09-24","arxiv_id":"2109.11778","n_code_links":0,"syntology":null},{"paper":null,"title":"What Vision-Language Models `See' when they See Scenes","date":"2021-09-15","arxiv_id":"2109.07301","n_code_links":0,"syntology":null},{"paper":"/paper/data-efficient-masked-language-modeling-for","title":"Data Efficient Masked Language Modeling for Vision and Language","date":"2021-09-05","arxiv_id":"2109.02040","n_code_links":1,"syntology":null},{"paper":"/paper/learning-to-read-maps-understanding-natural","title":"Learning to Read Maps: Understanding Natural Language Instructions from Unseen Maps","date":"2021-08-01","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/multi-stage-pre-training-over-simplified","title":"Multi-stage Pre-training over Simplified Multimodal Pre-training Models","date":"2021-07-22","arxiv_id":"2107.14596","n_code_links":1,"syntology":{"ran":3,"of":9,"unverified":6,"pointer_only":9}},{"paper":"/paper/passage-retrieval-for-outside-knowledge","title":"Passage Retrieval for Outside-Knowledge Visual Question Answering","date":"2021-05-09","arxiv_id":"2105.03938","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":2}},{"paper":null,"title":"Chop Chop BERT: Visual Question Answering by Chopping VisualBERT's Heads","date":"2021-04-30","arxiv_id":"2104.14741","n_code_links":0,"syntology":null},{"paper":null,"title":"Playing Lottery Tickets with Vision and Language","date":"2021-04-23","arxiv_id":"2104.11832","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/question-answering","name":"Question Answering","papers":18},{"task":"/task/visual-question-answering-1","name":"Visual Question Answering","papers":18},{"task":"/task/visual-question-answering","name":"Visual Question Answering (VQA)","papers":18},{"task":"/task/language-modeling","name":"Language Modeling","papers":5},{"task":"/task/language-modelling","name":"Language Modelling","papers":5},{"task":"/task/retrieval","name":"Retrieval","papers":5},{"task":"/task/sentence","name":"Sentence","papers":5},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":4},{"task":"/task/image-captioning","name":"Image Captioning","papers":4},{"task":"/task/image-text-matching","name":"Image-text matching","papers":4},{"task":"/task/text-matching","name":"Text Matching","papers":4},{"task":"/task/visual-reasoning","name":"Visual Reasoning","papers":4},{"task":"/task/image-text-retrieval","name":"Image-text Retrieval","papers":3},{"task":"/task/object","name":"Object","papers":3},{"task":"/task/object-localization","name":"Object Localization","papers":3},{"task":"/task/representation-learning","name":"Representation Learning","papers":3},{"task":"/task/text-retrieval","name":"Text Retrieval","papers":3},{"task":"/task/all","name":"All","papers":2},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":2},{"task":"/task/coreference-resolution","name":"Coreference Resolution","papers":2}],"tasks_shown":20,"n_tasks":44,"usage_by_year":[{"year":"2019","papers":1},{"year":"2020","papers":7},{"year":"2021","papers":13},{"year":"2022","papers":11},{"year":"2023","papers":6},{"year":"2024","papers":2}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/lxmert"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}