Papers › OmniVL:One Foundation Model for Image-Language and Video-Language Tasks

OmniVL:One Foundation Model for Image-Language and Video-Language Tasks

15 Sep 2022arXiv:2209.07526archive 2025-07-28

Junke Wang, Dongdong Chen, Zuxuan Wu, Chong Luo, Luowei Zhou, Yucheng Zhao, Yujia Xie, Ce Liu, Yu-Gang Jiang, Lu Yuan

This paper presents OmniVL, a new foundation model to support both image-language and video-language tasks using one universal architecture. It adopts a unified transformer-based visual encoder for both image and video inputs, and thus can perform joint image-language and video-language pretraining. We demonstrate, for the first time, such a paradigm benefits both image and video tasks, as opposed to the conventional one-directional transfer (e.g., use image-language to help video-language). To this end, we propose a decoupled joint pretraining of image-language and video-language to effectively decompose the vision-language modeling into spatial and temporal dimensions and obtain performance boost on both image and video tasks. Moreover, we introduce a novel unified vision-language contrastive (UniVLC) loss to leverage image-text, video-text, image-label (e.g., image classification), video-label (e.g., video action recognition) data together, so that both supervised and noisily supervised pretraining data are utilized as much as possible. Without incurring extra task-specific adaptors, OmniVL can simultaneously support visual only tasks (e.g., image classification, video action recognition), cross-modal alignment tasks (e.g., image/video-text retrieval), and multi-modal understanding and generation tasks (e.g., image/video question answering, captioning). We evaluate OmniVL on a wide range of downstream tasks and achieve state-of-the-art or competitive results with similar model size and data scale.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAction RecognitionCross-Modal RetrievalImage CaptioningImage ClassificationLanguage ModelingLanguage ModellingQuestion AnsweringRetrievalTemporal Action LocalizationText RetrievalVideo CaptioningVideo Question AnsweringVideo RetrievalVideo-Text RetrievalVisual Question Answering (VQA)Zero-Shot Video Retrievalcross-modal alignmentimage-classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Classification Kinetics-400 OmniVL Acc@1 79.1 #117 of 207 Archive leaderboard report
Action Classification Kinetics-400 OmniVL Acc@5 94.5 #117 of 207 Archive leaderboard report
Action Recognition Something-Something V2 OmniVL Top-1 Accuracy 62.5 #102 of 123 Archive leaderboard report
Action Recognition Something-Something V2 OmniVL Top-5 Accuracy 86.2 #102 of 123 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 OmniVL (14M) Image-to-text R@1 82.1 #7 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 OmniVL (14M) Image-to-text R@10 98.1 #7 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 OmniVL (14M) Image-to-text R@5 95.9 #7 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 OmniVL (14M) Text-to-image R@1 64.8 #7 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 OmniVL (14M) Text-to-image R@10 91.6 #7 of 36 Archive leaderboard report
Cross-Modal Retrieval COCO 2014 OmniVL (14M) Text-to-image R@5 86.1 #7 of 36 Archive leaderboard report
Cross-Modal Retrieval Flickr30k OmniVL (14M) Image-to-text R@1 97.3 #4 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k OmniVL (14M) Image-to-text R@10 100 #4 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k OmniVL (14M) Image-to-text R@5 99.9 #4 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k OmniVL (14M) Text-to-image R@1 87.9 #4 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k OmniVL (14M) Text-to-image R@10 99.1 #4 of 27 Archive leaderboard report
Cross-Modal Retrieval Flickr30k OmniVL (14M) Text-to-image R@5 97.8 #4 of 27 Archive leaderboard report
Image Captioning nocaps-val-in-domain OmniVL CIDEr 104.6 #9 of 11 Archive leaderboard report
Image Captioning nocaps-val-in-domain OmniVL Pre-train (#images) 14M #9 of 11 Archive leaderboard report
Image Captioning nocaps-val-in-domain OmniVL SPICE 15 #9 of 11 Archive leaderboard report
Image Captioning nocaps-val-near-domain OmniVL CIDEr 108.3 #8 of 10 Archive leaderboard report
Image Captioning nocaps-val-near-domain OmniVL Pre-train (#images) 14M #8 of 10 Archive leaderboard report
Image Captioning nocaps-val-near-domain OmniVL SPICE 14.9 #8 of 10 Archive leaderboard report
Image Captioning nocaps-val-out-domain OmniVL CIDEr 106.3 #8 of 10 Archive leaderboard report
Image Captioning nocaps-val-out-domain OmniVL Pretrain (#images) 14M #8 of 10 Archive leaderboard report
Image Captioning nocaps-val-out-domain OmniVL SPICE 14.2 #8 of 10 Archive leaderboard report
Image Captioning nocaps-val-overall OmniVL CIDEr 107.5 #8 of 11 Archive leaderboard report
Image Captioning nocaps-val-overall OmniVL Pretrain (#images) 14M #8 of 11 Archive leaderboard report
Image Captioning nocaps-val-overall OmniVL SPICE 14.7 #8 of 11 Archive leaderboard report
Video Captioning YouCook2 OmniVL BLEU-3 12.87 #11 of 14 Archive leaderboard report
Video Captioning YouCook2 OmniVL BLEU-4 8.72 #11 of 14 Archive leaderboard report
Video Captioning YouCook2 OmniVL CIDEr 1.16 #11 of 14 Archive leaderboard report
Video Captioning YouCook2 OmniVL METEOR 14.83 #11 of 14 Archive leaderboard report
Video Captioning YouCook2 OmniVL ROUGE-L 36.09 #11 of 14 Archive leaderboard report
Video Retrieval DiDeMo OmniVL text-to-video R@1 52.4 #21 of 40 Archive leaderboard report
Video Retrieval DiDeMo OmniVL text-to-video R@10 85.4 #21 of 40 Archive leaderboard report
Video Retrieval DiDeMo OmniVL text-to-video R@5 79.5 #21 of 40 Archive leaderboard report
Video Retrieval MSR-VTT OmniVL text-to-video R@1 47.8 #13 of 40 Archive leaderboard report
Video Retrieval MSR-VTT OmniVL text-to-video R@10 83.8 #13 of 40 Archive leaderboard report
Video Retrieval MSR-VTT OmniVL text-to-video R@5 74.2 #13 of 40 Archive leaderboard report
Visual Question Answering (VQA) MSRVTT-QA OmniVL Accuracy 0.441 #18 of 34 Archive leaderboard report
Visual Question Answering (VQA) MSVD-QA OmniVL Accuracy 0.510 #21 of 36 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo OmniVL text-to-video R@1 33.3 #15 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo OmniVL text-to-video R@10 68.5 #15 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo OmniVL text-to-video R@5 58.7 #15 of 26 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT OmniVL text-to-video R@1 34.6 #18 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT OmniVL text-to-video R@10 66.6 #18 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT OmniVL text-to-video R@5 58.4 #18 of 41 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections