{"url":"/method/unified-vlp","slug":"unified-vlp","name":"Unified VLP","full_name":"Unified VLP","full_name_withheld":false,"description_markdown":"Unified VLP is unified encoder-decoder model for general vision-language pre-training. The models uses a shared multi-layer transformers network for both encoding and decoding. The model is pre-trained on large amount of image-text pairs using the unsupervised learning objectives of two tasks: bidirectional and sequence-to-sequence (seq2seq) masked vision-language prediction. Model architecture for pre-training. For pre-training , the input comprises of image input, sentence input, and three special tokens ([CLS], [SEP], [STOP]). The image is processed as $N$ Region of Interests (RoIs) and region features are extracted. The sentence is tokenized and masked with [MASK] tokens for the later masked language modeling task. The model consists of 12 layers of Transformer blocks, each having a masked self-attention layer and feed-forward module, where the self-attention mask controls what input context the prediction conditions on. Two self-attention masks are implemented depending on whether the objective is bidirectional or seq2seq. The model is fine-tuned for image captioning and visual question answering.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/1909.11059v3","title":"Unified Vision-Language Pre-Training for Image Captioning and VQA","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision and Language Pre-Trained Models","url":"/methods/category/vision-and-language-pre-trained-models","pwc_aliases":[]}],"n_papers_tagged":2,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations","date":"2025-03-02","arxiv_id":"2503.01019","n_code_links":0,"syntology":null},{"paper":"/paper/unified-vision-language-pre-training-for","title":"Unified Vision-Language Pre-Training for Image Captioning and VQA","date":"2019-09-24","arxiv_id":"1909.11059","n_code_links":3,"syntology":{"ran":5,"of":14,"unverified":9,"pointer_only":13}}],"papers_shown":2,"tasks":[{"task":"/task/text-generation","name":"Text Generation","papers":2},{"task":"/task/decoder","name":"Decoder","papers":1},{"task":"/task/image-captioning","name":"Image Captioning","papers":1},{"task":"/task/image-classification","name":"Image Classification","papers":1},{"task":"/task/image-generation","name":"Image Generation","papers":1},{"task":"/task/image-text-retrieval","name":"Image-text Retrieval","papers":1},{"task":"/task/image-text-matching","name":"Image-text matching","papers":1},{"task":"/task/medical-report-generation","name":"Medical Report Generation","papers":1},{"task":"/task/quantization","name":"Quantization","papers":1},{"task":"/task/question-answering","name":"Question Answering","papers":1},{"task":"/task/text-matching","name":"Text Matching","papers":1},{"task":"/task/text-retrieval","name":"Text Retrieval","papers":1},{"task":"/task/visual-question-answering-1","name":"Visual Question Answering","papers":1},{"task":"/task/visual-question-answering","name":"Visual Question Answering (VQA)","papers":1},{"task":"/task/zero-shot-image-classification","name":"Zero-Shot Image Classification","papers":1},{"task":"/task/image-classification","name":"image-classification","papers":1}],"tasks_shown":16,"n_tasks":16,"usage_by_year":[{"year":"2019","papers":1},{"year":"2025","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/unified-vlp"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}