{"url":"/method/vlmo","slug":"vlmo","name":"VLMo","full_name":"Vision-Language pretrained Model","full_name_withheld":false,"description_markdown":"VLMo is a unified vision-language pre-trained model that jointly learns a dual encoder and a fusion encoder with a modular Transformer network. A Mixture-of-Modality-Experts (MOME) transformer is introduced to encode different modalities which helps it to capture modality-specific information by modality experts, and align content of different modalities by the self-attention module shared across modalities. The model parameters are shared across image-text contrastive learning, masked language modeling, and image-text matching tasks. During fine-tuning, the flexible modeling allows for VLMO to be used as either a dual encoder (i.e., separately encode images and text for retrieval tasks) or a fusion encoder (i.e., jointly encode image-text pairs for better interaction across modalities) Stage-wise pretraining on image-only and text-only data improved the vision-language pre-trained model. The model can be used for classification tasks and fine-tuned as a dual encoder for retrieval tasks.","description_state":"present","introduced_year":null,"introduced_by":{"title":"VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts","paper":"/paper/vlmo-unified-vision-language-pre-training","first_author":"Hangbo Bao","n_authors":8,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/vlmo-unified-vision-language-pre-training"},"source":{"url":"https://arxiv.org/abs/2111.02358v2","title":"VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision and Language Pre-Trained Models","url":"/methods/category/vision-and-language-pre-trained-models","pwc_aliases":[]}],"n_papers_tagged":1,"archive_num_papers":1,"papers_newest_first":[{"paper":"/paper/vlmo-unified-vision-language-pre-training","title":"VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts","date":"2021-11-03","arxiv_id":"2111.02358","n_code_links":2,"syntology":null}],"papers_shown":1,"tasks":[{"task":"/task/image-retrieval","name":"Image Retrieval","papers":1},{"task":"/task/image-text-retrieval","name":"Image-text Retrieval","papers":1},{"task":"/task/retrieval","name":"Retrieval","papers":1},{"task":"/task/text-retrieval","name":"Text Retrieval","papers":1},{"task":"/task/visual-question-answering","name":"Visual Question Answering (VQA)","papers":1},{"task":"/task/visual-reasoning","name":"Visual Reasoning","papers":1}],"tasks_shown":6,"n_tasks":6,"usage_by_year":[{"year":"2021","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/vlmo"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}