Methods › Computer Vision › Vision and Language Pre-Trained Models › VisualBERT

VisualBERT

25 papers tagged archive 2025-07-28

Introduced by Liunian Harold Li et al. in VisualBERT: A Simple and Performant Baseline for Vision and Language

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

VisualBERT aims to reuse self-attention to implicitly align elements of the input text and regions in the input image. Visual embeddings are used to model images where the representations are represented by a bounding region in an image obtained from an object detector. These visual embeddings are constructed by summing three embeddings: 1) visual feature representation, 2) a segment embedding indicate whether it is an image embedding, and 3) position embedding. Essentially, image regions and language are combined with a Transformer to allow self-attention to discover implicit alignments between language and vision. VisualBERT is trained using COCO, which consists of images paired with captions. It is pre-trained using two objectives: masked language modeling objective and sentence-image prediction task. It can then be fine-tuned on different downstream tasks.

PaperSource

Papers archive 2025-07-28

25 shown of 25, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 45 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Visual Question Answering (VQA)9
Question Answering7
Visual Question Answering7
Language Modeling4
Language Modelling4
Visual Reasoning4
Image Captioning3
Representation Learning3
Multimodal Deep Learning2
Object2
Visual Commonsense Reasoning2
Visual Entailment2
All1
Conditional Image Generation1
Cultural Vocal Bursts Intensity Prediction1
Ensemble Learning1
Factual Visual Question Answering1
Fairness1
Image-text Retrieval1
Image-text matching1

Usage over time archive 2025-07-28

Papers per year tagged with VisualBERT: 2019 to 2025, peak 10 10 0 2019: 1 paper 2019 2020: 2 papers 2020 2021: 10 papers 2021 2022: 6 papers 2022 2023: 2 papers 2023 2024: 3 papers 2024 2025: 1 paper 2025
Papers per year the archive tags with this method, by the paper's archive date (25 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Vision and Language Pre-Trained Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections