Methods › Computer Vision › Vision and Language Pre-Trained Models › LXMERT
Learning Cross-Modality Encoder Representations from Transformers
LXMERT
Introduced by Hao Tan et al. in LXMERT: Learning Cross-Modality Encoder Representations from Transformers
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
LXMERT is a model for learning vision-and-language cross-modality representations. It consists of a Transformer model that consists three encoders: object relationship encoder, a language encoder, and a cross-modality encoder. The model takes two inputs: image with its related sentence. The images are represented as a sequence of objects, whereas each sentence is represented as sequence of words. By combining the self-attention and cross-attention layers the model is able to generated language representation, image representations, and cross-modality representations from the input. The model is pre-trained with image-sentence pairs via five pre-training tasks: masked language modeling, masked object prediction, cross-modality matching, and image questions answering. These tasks help the model to learn both intra-modality and cross-modality relationships.
Papers archive 2025-07-28
30 shown of 40, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Optimizing Visual Question Answering Models for Driving: Bridging the Gap Between Human and Machine Attention Patterns 13 Jun 2024 · 0 repositories · arXiv:2406.09203
-
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking 29 Jan 2024 · 1 repository · arXiv:2401.16575
-
LXMERT Model Compression for Visual Question Answering 23 Oct 2023 · 2 repositories · arXiv:2310.15325
-
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models 18 Aug 2023 · 0 repositories · arXiv:2308.09778
-
Towards a performance analysis on pre-trained Visual Question Answering models for autonomous driving 18 Jul 2023 · 1 repository · arXiv:2307.09329
-
An Empirical Study on the Language Modal in Visual Question Answering 17 May 2023 · 0 repositories · arXiv:2305.10143
-
Probing the Role of Positional Information in Vision-Language Models 17 May 2023 · 0 repositories · arXiv:2305.10046
-
Controlling for Stereotypes in Multimodal Language Model Evaluation 3 Feb 2023 · 0 repositories · arXiv:2302.01582
-
MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks 15 Dec 2022 · 1 repository · arXiv:2212.08158Syntology ran 1 of 10 samples · 9 unverified
-
Enhancing Self-Consistency and Performance of Pre-Trained Language Models through Natural Language Inference 21 Nov 2022 · 0 repositories · arXiv:2211.11875
-
Compressing And Debiasing Vision-Language Pre-Trained Models for Visual Question Answering 26 Oct 2022 · 1 repository · arXiv:2210.14558
-
Probing Cross-modal Semantics Alignment Capability from the Textual Perspective 18 Oct 2022 · 0 repositories · arXiv:2210.09550
-
Generative Bias for Robust Visual Question Answering 1 Aug 2022 · 1 repository · arXiv:2208.00690Syntology ran 0 of 2 samples · 2 unverified
-
VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations 1 Jul 2022 · 1 repository · arXiv:2207.00221
-
Visual Spatial Reasoning 30 Apr 2022 · 4 repositories · arXiv:2205.00363Syntology ran 5 of 10 samples · 5 unverified
-
Visio-Linguistic Brain Encoding 18 Apr 2022 · 0 repositories · arXiv:2204.08261
-
SwapMix: Diagnosing and Regularizing the Over-Reliance on Visual Context in Visual Question Answering 5 Apr 2022 · 1 repository · arXiv:2204.02285
-
Exploring Multi-Modal Representations for Ambiguity Detection & Coreference Resolution in the SIMMC 2.0 Challenge 25 Feb 2022 · 2 repositories · arXiv:2202.12645
-
Probing the Role of Positional Information in Vision-Language Models 16 Jan 2022 · 0 repositories
-
Multimodal Learning: Are Captions All You Need? 16 Nov 2021 · 0 repositories
-
Region under Discussion for visual dialog 1 Nov 2021 · 0 repositories
-
Seeing things or seeing scenes: Investigating the capabilities of V&L models to align scene descriptions to images 16 Oct 2021 · 0 repositories
-
Dense Contrastive Visual-Linguistic Pretraining 24 Sep 2021 · 0 repositories · arXiv:2109.11778
-
What Vision-Language Models `See' when they See Scenes 15 Sep 2021 · 0 repositories · arXiv:2109.07301
-
Data Efficient Masked Language Modeling for Vision and Language 5 Sep 2021 · 1 repository · arXiv:2109.02040
-
Learning to Read Maps: Understanding Natural Language Instructions from Unseen Maps 1 Aug 2021 · 1 repository
-
Multi-stage Pre-training over Simplified Multimodal Pre-training Models 22 Jul 2021 · 1 repository · arXiv:2107.14596Syntology ran 3 of 9 samples · 6 unverified · 9 pointer-only (licence)
-
Passage Retrieval for Outside-Knowledge Visual Question Answering 9 May 2021 · 1 repository · arXiv:2105.03938Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)
-
Chop Chop BERT: Visual Question Answering by Chopping VisualBERT's Heads 30 Apr 2021 · 0 repositories · arXiv:2104.14741
-
Playing Lottery Tickets with Vision and Language 23 Apr 2021 · 0 repositories · arXiv:2104.11832
Tasks archive 2025-07-28
20 shown of 44 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections