Methods › Computer Vision › Vision and Language Pre-Trained Models › VisualBERT
VisualBERT
Introduced by Liunian Harold Li et al. in VisualBERT: A Simple and Performant Baseline for Vision and Language
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
VisualBERT aims to reuse self-attention to implicitly align elements of the input text and regions in the input image. Visual embeddings are used to model images where the representations are represented by a bounding region in an image obtained from an object detector. These visual embeddings are constructed by summing three embeddings: 1) visual feature representation, 2) a segment embedding indicate whether it is an image embedding, and 3) position embedding. Essentially, image regions and language are combined with a Transformer to allow self-attention to discover implicit alignments between language and vision. VisualBERT is trained using COCO, which consists of images paired with captions. It is pre-trained using two objectives: masked language modeling objective and sentence-image prediction task. It can then be fine-tuned on different downstream tasks.
Papers archive 2025-07-28
25 shown of 25, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Visual Question Answering on Multiple Remote Sensing Image Modalities 21 May 2025 · 0 repositories · arXiv:2505.15401
-
Seeing Through VisualBERT: A Causal Adventure on Memetic Landscapes 17 Oct 2024 · 0 repositories · arXiv:2410.13488
-
OSPC: Detecting Harmful Memes with Large Language Model as a Catalyst 14 Jun 2024 · 0 repositories · arXiv:2406.09779
-
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking 29 Jan 2024 · 1 repository · arXiv:2401.16575
-
A Review of Vision-Language Models and their Performance on the Hateful Memes Challenge 9 May 2023 · 1 repository · arXiv:2305.06159
-
Controlling for Stereotypes in Multimodal Language Model Evaluation 3 Feb 2023 · 0 repositories · arXiv:2302.01582
-
A survey on knowledge-enhanced multimodal learning 19 Nov 2022 · 0 repositories · arXiv:2211.12328
-
Transfer Learning with Joint Fine-Tuning for Multimodal Sentiment Analysis 11 Oct 2022 · 1 repository · arXiv:2210.05790
-
Multi-Modal Fusion Transformer for Visual Question Answering in Remote Sensing 10 Oct 2022 · 0 repositories · arXiv:2210.04510
-
Surgical-VQA: Visual Question Answering in Surgical Scenes using Transformer 22 Jun 2022 · 3 repositories · arXiv:2206.11053Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)
-
Visual Spatial Reasoning 30 Apr 2022 · 4 repositories · arXiv:2205.00363Syntology ran 5 of 10 samples · 5 unverified
-
Visio-Linguistic Brain Encoding 18 Apr 2022 · 0 repositories · arXiv:2204.08261
-
Hateful Memes Challenge: An Enhanced Multimodal Framework 20 Dec 2021 · 1 repository · arXiv:2112.11244
-
Multimodal Learning: Are Captions All You Need? 16 Nov 2021 · 0 repositories
-
Seeing things or seeing scenes: Investigating the capabilities of V&L models to align scene descriptions to images 16 Oct 2021 · 0 repositories
-
Understanding of Emotion Perception from Art 13 Oct 2021 · 0 repositories · arXiv:2110.06486
-
MARMOT: A Deep Learning Framework for Constructing Multimodal Representations for Vision-and-Language Tasks 23 Sep 2021 · 1 repository · arXiv:2109.11526
-
What Vision-Language Models `See' when they See Scenes 15 Sep 2021 · 0 repositories · arXiv:2109.07301
-
Broaden the Vision: Geo-Diverse Visual Commonsense Reasoning 14 Sep 2021 · 1 repository · arXiv:2109.06860Syntology ran 0 of 9 samples · 9 unverified
-
BERTHop: An Effective Vision-and-Language Model for Chest X-ray Disease Diagnosis 10 Aug 2021 · 1 repository · arXiv:2108.04938
-
Chop Chop BERT: Visual Question Answering by Chopping VisualBERT's Heads 30 Apr 2021 · 0 repositories · arXiv:2104.14741
-
Analysis on Image Set Visual Question Answering 31 Mar 2021 · 0 repositories · arXiv:2104.00107
-
Detecting Hate Speech in Memes Using Multimodal Deep Learning Approaches: Prize-winning solution to Hateful Memes Challenge 23 Dec 2020 · 1 repository · arXiv:2012.12975
-
A Comparison of Pre-trained Vision-and-Language Models for Multimodal Representation Learning across Medical Images and Reports 3 Sep 2020 · 1 repository · arXiv:2009.01523
-
VisualBERT: A Simple and Performant Baseline for Vision and Language 9 Aug 2019 · 10 repositories · arXiv:1908.03557Syntology ran 4 of 9 samples · 5 unverified · 6 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 45 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections