Methods › Computer Vision › Vision and Language Pre-Trained Models › ViLBERT
Vision-and-Language BERT
ViLBERT
Introduced by Jiasen Lu et al. in ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Vision-and-Language BERT (ViLBERT) is a BERT-based model for learning task-agnostic joint representations of image content and natural language. ViLBERT extend the popular BERT architecture to a multi-modal two-stream model, processing both visual and textual inputs in separate streams that interact through co-attentional transformer layers.
Papers archive 2025-07-28
30 shown of 30, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
A Multimodal Fusion Network For Student Emotion Recognition Based on Transformer and Tensor Product 13 Mar 2024 · 0 repositories · arXiv:2403.08511
-
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking 29 Jan 2024 · 1 repository · arXiv:2401.16575
-
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models 23 Oct 2023 · 1 repository · arXiv:2310.15061Syntology ran 11 of 12 samples · 1 unverified
-
Towards a performance analysis on pre-trained Visual Question Answering models for autonomous driving 18 Jul 2023 · 1 repository · arXiv:2307.09329
-
Switch-BERT: Learning to Model Multimodal Interactions by Switching Attention and Input 25 Jun 2023 · 0 repositories · arXiv:2306.14182
-
Weakly Supervised Visual Question Answer Generation 11 Jun 2023 · 0 repositories · arXiv:2306.06622
-
Unified Multimodal Model with Unlikelihood Training for Visual Dialog 23 Nov 2022 · 1 repository · arXiv:2211.13235
-
A survey on knowledge-enhanced multimodal learning 19 Nov 2022 · 0 repositories · arXiv:2211.12328
-
Probing Cross-modal Semantics Alignment Capability from the Textual Perspective 18 Oct 2022 · 0 repositories · arXiv:2210.09550
-
Image Retrieval from Contextual Descriptions 29 Mar 2022 · 1 repository · arXiv:2203.15867Syntology ran 2 of 2 samples · 0 unverified
-
TriBERT: Human-centric Audio-visual Representation Learning 1 Dec 2021 · 1 repository
-
Image Retrieval from Contextual Descriptions 16 Nov 2021 · 0 repositories
-
Multimodal Learning: Are Captions All You Need? 16 Nov 2021 · 0 repositories
-
TriBERT: Full-body Human-centric Audio-visual Representation Learning for Visual Sound Separation 26 Oct 2021 · 1 repository · arXiv:2110.13412Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)
-
Broaden the Vision: Geo-Diverse Visual Commonsense Reasoning 14 Sep 2021 · 1 repository · arXiv:2109.06860Syntology ran 0 of 9 samples · 9 unverified
-
Enhance Multimodal Model Performance with Data Augmentation: Facebook Hateful Meme Challenge Solution 25 May 2021 · 1 repository · arXiv:2105.13132
-
Chop Chop BERT: Visual Question Answering by Chopping VisualBERT's Heads 30 Apr 2021 · 0 repositories · arXiv:2104.14741
-
Playing Lottery Tickets with Vision and Language 23 Apr 2021 · 0 repositories · arXiv:2104.11832
-
On the Role of Images for Analyzing Claims in Social Media 17 Mar 2021 · 1 repository · arXiv:2103.09602
-
Seeing past words: Testing the cross-modal capabilities of pretrained V&L models on counting tasks 22 Dec 2020 · 0 repositories · arXiv:2012.12352
-
A Closer Look at the Robustness of Vision-and-Language Pre-trained Models 15 Dec 2020 · 0 repositories · arXiv:2012.08673
-
A Multi-Modal Method for Satire Detection using Textual and Visual Cues 13 Oct 2020 · 1 repository · arXiv:2010.06671
-
X-LXMERT: Paint, Caption and Answer Questions with Multi-Modal Transformers 23 Sep 2020 · 2 repositories · arXiv:2009.11278Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)
-
Contrastive Visual-Linguistic Pretraining 26 Jul 2020 · 0 repositories · arXiv:2007.13135
-
What Does BERT with Vision Look At? 1 Jul 2020 · 0 repositories
-
Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models 15 May 2020 · 0 repositories · arXiv:2005.07310
-
Words aren't enough, their order matters: On the Robustness of Grounding Visual Referring Expressions 4 May 2020 · 1 repository · arXiv:2005.01655
-
Generating Rationales in Visual Question Answering 4 Apr 2020 · 0 repositories · arXiv:2004.02032
-
Large-scale Pretraining for Visual Dialog: A Simple State-of-the-Art Baseline 5 Dec 2019 · 2 repositories · arXiv:1912.02379Syntology ran 3 of 6 samples · 3 unverified
-
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks 6 Aug 2019 · 11 repositories · arXiv:1908.02265Syntology ran 10 of 34 samples · 24 unverified · 34 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 64 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections