Methods › Computer Vision › Vision and Language Pre-Trained Models › VLMo
Vision-Language pretrained Model
VLMo
Introduced by Hangbo Bao et al. in VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
VLMo is a unified vision-language pre-trained model that jointly learns a dual encoder and a fusion encoder with a modular Transformer network. A Mixture-of-Modality-Experts (MOME) transformer is introduced to encode different modalities which helps it to capture modality-specific information by modality experts, and align content of different modalities by the self-attention module shared across modalities. The model parameters are shared across image-text contrastive learning, masked language modeling, and image-text matching tasks. During fine-tuning, the flexible modeling allows for VLMO to be used as either a dual encoder (i.e., separately encode images and text for retrieval tasks) or a fusion encoder (i.e., jointly encode image-text pairs for better interaction across modalities) Stage-wise pretraining on image-only and text-only data improved the vision-language pre-trained model. The model can be used for classification tasks and fine-tuned as a dual encoder for retrieval tasks.
Papers archive 2025-07-28
1 shown of 1, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts 3 Nov 2021 · 2 repositories · arXiv:2111.02358
Tasks archive 2025-07-28
6 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| Image Retrieval | 1 |
| Image-text Retrieval | 1 |
| Retrieval | 1 |
| Text Retrieval | 1 |
| Visual Question Answering (VQA) | 1 |
| Visual Reasoning | 1 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections