Methods › Computer Vision › Vision and Language Pre-Trained Models

Vision and Language Pre-Trained Models

30 methods 8,661 papers tagged archive 2025-07-28

Involves models that adapt pre-training to the field of Vision-and-Language (V-L) learning and improve the performance on downstream tasks like visual question answering and visual captioning.

According to Du et al. (2022), information coming from the different modalities can be encoded in three ways: fusion encoder, dual encoder, and a combination of both.

References:

Methods

All 30 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.

ALIGN – 5,527
CLIP Contrastive Language-Image Pre-training – 3,094
BLIP BLIP: Bootstrapping Language-Image Pre-training – 93
LXMERT Learning Cross-Modality Encoder Representations from Transformers – 40
OSCAR – 36
OFA – 32
ViLBERT Vision-and-Language BERT – 30
VisualBERT – 25
ALBEF – 17
Florence – 12
FLAVA – 11
ViLT Vision-and-Language Transformer – 6
Visual Parsing – 6
AltCLIP – 5
InternVideo InternVideo: General Video Foundation Models via Generative and Discriminative Learning – 5
PLIP Pathology Language and Image Pre-Training – 5
SOHO – 5
VL-T5 – 5
UNIMO – 4
VL-BERT Visual-Linguistic BERT – 4
WenLan – 4
SimVLM Simple Visual Language Model – 3
FashionCLIP – 2
Kaleido-BERT – 2
OneR One Representation – 2
Pixel-BERT – 2
Unified VLP – 2
InterBERT – 1
VLMo Vision-Language pretrained Model – 1
XGPT – 1