Methods › Computer Vision › Vision and Language Pre-Trained Models › VL-BERT
Visual-Linguistic BERT
VL-BERT
Introduced by Weijie Su et al. in VL-BERT: Pre-training of Generic Visual-Linguistic Representations
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
VL-BERT is pre-trained on a large-scale image-captions dataset together with text-only corpus. The input to the model are either words from the input sentences or regions-of-interest (RoI) from input images. It can be fine-tuned to fit most visual-linguistic downstream tasks. Its backbone is a multi-layer bidirectional Transformer encoder, modified to accommodate visual contents, and new type of visual feature embedding to the input feature embeddings. VL-BERT takes both visual and linguistic elements as input, represented as RoIs in images and subwords in input sentences. Four different types of embeddings are used to represent each input: token embedding, visual feature embedding, segment embedding, and sequence position embedding. VL-BERT is pre-trained using Conceptual Captions and text-only datasets. Two pre-training tasks are used: masked language modeling with visual clues, and masked RoI classification with linguistic clues.
Papers archive 2025-07-28
4 shown of 4, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Self-Training Vision Language BERTs with a Unified Conditional Model 6 Jan 2022 · 0 repositories · arXiv:2201.02010
-
BERTGEN: Multi-task Generation through BERT 7 Jun 2021 · 1 repository · arXiv:2106.03484
-
Worst of Both Worlds: Biases Compound in Pre-trained Vision-and-Language Models 18 Apr 2021 · 0 repositories · arXiv:2104.08666
-
VL-BERT: Pre-training of Generic Visual-Linguistic Representations 22 Aug 2019 · 3 repositories · arXiv:1908.08530Syntology ran 0 of 1 samples · 1 unverified
Tasks archive 2025-07-28
15 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections