Papers › VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending

VLAB: Enhancing Video Language Pre-training by Feature Adapting and Blending

22 May 2023arXiv:2305.13167archive 2025-07-28

Xingjian He, Sihan Chen, Fan Ma, Zhicheng Huang, Xiaojie Jin, Zikang Liu, Dongmei Fu, Yi Yang, Jing Liu, Jiashi Feng

Large-scale image-text contrastive pre-training models, such as CLIP, have been demonstrated to effectively learn high-quality multimodal representations. However, there is limited research on learning video-text representations for general video multimodal tasks based on these powerful features. Towards this goal, we propose a novel video-text pre-training method dubbed VLAB: Video Language pre-training by feature Adapting and Blending, which transfers CLIP representations to video pre-training tasks and develops unified video multimodal models for a wide range of video-text tasks. Specifically, VLAB is founded on two key strategies: feature adapting and feature blending. In the former, we introduce a new video adapter module to address CLIP's deficiency in modeling temporal information and extend the model's capability to encompass both contrastive and generative tasks. In the latter, we propose an end-to-end training method that further enhances the model's performance by exploiting the complementarity of image and video features. We validate the effectiveness and versatility of VLAB through extensive experiments on highly competitive video multimodal tasks, including video text retrieval, video captioning, and video question answering. Remarkably, VLAB outperforms competing methods significantly and sets new records in video question answering on MSRVTT, MSVD, and TGIF datasets. It achieves an accuracy of 49.6, 61.0, and 79.0, respectively. Codes and models will be released.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Question AnsweringRetrievalTGIF-FrameText RetrievalVideo CaptioningVideo Question AnsweringVideo RetrievalVideo-Text RetrievalVisual Question Answering (VQA)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Captioning MSR-VTT VLAB BLEU-4 54.6 #4 of 24 Archive leaderboard report
Video Captioning MSR-VTT VLAB CIDEr 74.9 #4 of 24 Archive leaderboard report
Video Captioning MSR-VTT VLAB METEOR 33.4 #4 of 24 Archive leaderboard report
Video Captioning MSR-VTT VLAB ROUGE-L 68.3 #4 of 24 Archive leaderboard report
Video Captioning MSVD VLAB BLEU-4 79.3 #2 of 16 Archive leaderboard report
Video Captioning MSVD VLAB CIDEr 179.8 #2 of 16 Archive leaderboard report
Video Captioning MSVD VLAB METEOR 51.2 #2 of 16 Archive leaderboard report
Video Captioning MSVD VLAB ROUGE-L 87.9 #2 of 16 Archive leaderboard report
Video Retrieval DiDeMo VLAB text-to-video R@1 56.8 #12 of 40 Archive leaderboard report
Video Retrieval DiDeMo VLAB text-to-video R@10 88.7 #12 of 40 Archive leaderboard report
Video Retrieval DiDeMo VLAB text-to-video R@5 81.6 #12 of 40 Archive leaderboard report
Video Retrieval MSR-VTT VLAB text-to-video R@1 55.1 #9 of 40 Archive leaderboard report
Video Retrieval MSR-VTT VLAB text-to-video R@10 87.6 #9 of 40 Archive leaderboard report
Video Retrieval MSR-VTT VLAB text-to-video R@5 78.8 #9 of 40 Archive leaderboard report
Video Retrieval MSVD VLAB text-to-video R@1 57.5 #6 of 24 Archive leaderboard report
Video Retrieval MSVD VLAB text-to-video R@10 89.9 #6 of 24 Archive leaderboard report
Video Retrieval MSVD VLAB text-to-video R@5 83.6 #6 of 24 Archive leaderboard report
Visual Question Answering (VQA) MSRVTT-QA VLAB Accuracy 0.496 #1 of 34 Archive leaderboard report
Visual Question Answering (VQA) MSVD-QA VLAB Accuracy 0.61 #1 of 36 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdapterCLIP

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections