Methods › Computer Vision › Vision and Language Pre-Trained Models › VL-BERT

Visual-Linguistic BERT

VL-BERT

4 papers tagged archive 2025-07-28

Introduced by Weijie Su et al. in VL-BERT: Pre-training of Generic Visual-Linguistic Representations

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

VL-BERT is pre-trained on a large-scale image-captions dataset together with text-only corpus. The input to the model are either words from the input sentences or regions-of-interest (RoI) from input images. It can be fine-tuned to fit most visual-linguistic downstream tasks. Its backbone is a multi-layer bidirectional Transformer encoder, modified to accommodate visual contents, and new type of visual feature embedding to the input feature embeddings. VL-BERT takes both visual and linguistic elements as input, represented as RoIs in images and subwords in input sentences. Four different types of embeddings are used to represent each input: token embedding, visual feature embedding, segment embedding, and sequence position embedding. VL-BERT is pre-trained using Conceptual Captions and text-only datasets. Two pre-training tasks are used: masked language modeling with visual clues, and masked RoI classification with linguistic clues.

PaperSource

Papers archive 2025-07-28

4 shown of 4, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

15 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Decoder1
Image Captioning1
Image-text matching1
Language Modelling1
Machine Translation1
Multimodal Machine Translation1
Question Answering1
Referring Expression1
Referring Expression Comprehension1
Sentence1
Text Generation1
Translation1
Visual Commonsense Reasoning1
Visual Question Answering1
Visual Question Answering (VQA)1

Usage over time archive 2025-07-28

Papers per year tagged with VL-BERT: 2019 to 2022, peak 2 2 0 2019: 1 paper 2019 2020: 0 papers 2020 2021: 2 papers 2021 2022: 1 paper 2022
Papers per year the archive tags with this method, by the paper's archive date (4 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Vision and Language Pre-Trained Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections