Papers › Clover: Towards A Unified Video-Language Alignment and Fusion Model

Clover: Towards A Unified Video-Language Alignment and Fusion Model

16 Jul 2022CVPR 2023 1arXiv:2207.07885archive 2025-07-28

Jingjia Huang, Yinan Li, Jiashi Feng, Xinglong Wu, Xiaoshuai Sun, Rongrong Ji

Building a universal Video-Language model for solving various video understanding tasks (\emph{e.g.}, text-video retrieval, video question answering) is an open challenge to the machine learning field. Towards this goal, most recent works build the model by stacking uni-modal and cross-modal feature encoders and train it with pair-wise contrastive pre-text tasks. Though offering attractive generality, the resulted models have to compromise between efficiency and performance. They mostly adopt different architectures to deal with different downstream tasks. We find this is because the pair-wise training cannot well \emph{align} and \emph{fuse} features from different modalities. We then introduce \textbf{Clover}\textemdash a Correlated Video-Language pre-training method\textemdash towards a universal Video-Language model for solving multiple video understanding tasks with neither performance nor efficiency compromise. It improves cross-modal feature alignment and fusion via a novel tri-modal alignment pre-training task. Additionally, we propose to enhance the tri-modal alignment via incorporating learning from semantic masked samples and a new pair-wise ranking loss. Clover establishes new state-of-the-arts on multiple downstream tasks, including three retrieval tasks for both zero-shot and fine-tuning settings, and eight video question answering tasks. Codes and pre-trained models will be released at \url{https://github.com/LeeYN-43/Clover}.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

leeyn-43/clover officialmentioned in papermentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Language ModelingLanguage ModellingQuestion AnsweringRetrievalTGIF-ActionTGIF-FrameTGIF-TransitionText to Video RetrievalVideo Question AnsweringVideo RetrievalVideo UnderstandingVisual Question Answering (VQA)Zero-Shot Video Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Question Answering LSMDC-FiB Clover Accuracy 54.1 #1 of 1 Archive leaderboard report
Video Question Answering LSMDC-MC Clover Accuracy 83.7 #2 of 2 Archive leaderboard report
Video Question Answering MSRVTT-MC Clover Accuracy 95.2 #4 of 7 Archive leaderboard report
Video Retrieval DiDeMo Clover text-to-video Median Rank 1 #24 of 40 Archive leaderboard report
Video Retrieval DiDeMo Clover text-to-video R@1 50.1 #24 of 40 Archive leaderboard report
Video Retrieval DiDeMo Clover text-to-video R@10 85.6 #24 of 40 Archive leaderboard report
Video Retrieval DiDeMo Clover text-to-video R@5 76.7 #24 of 40 Archive leaderboard report
Video Retrieval LSMDC Clover text-to-video Median Rank 8 #18 of 38 Archive leaderboard report
Video Retrieval LSMDC Clover text-to-video R@1 24.8 #18 of 38 Archive leaderboard report
Video Retrieval LSMDC Clover text-to-video R@10 54.5 #18 of 38 Archive leaderboard report
Video Retrieval LSMDC Clover text-to-video R@5 44 #18 of 38 Archive leaderboard report
Video Retrieval MSR-VTT-1kA Clover text-to-video Median Rank 2 #39 of 63 Archive leaderboard report
Video Retrieval MSR-VTT-1kA Clover text-to-video R@1 40.5 #39 of 63 Archive leaderboard report
Video Retrieval MSR-VTT-1kA Clover text-to-video R@10 79.4 #39 of 63 Archive leaderboard report
Video Retrieval MSR-VTT-1kA Clover text-to-video R@5 69.8 #39 of 63 Archive leaderboard report
Visual Question Answering (VQA) MSRVTT-QA Clover Accuracy 0.441 #19 of 34 Archive leaderboard report
Visual Question Answering (VQA) MSVD-QA Clover Accuracy 0.524 #19 of 36 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo Clover text-to-video Median Rank 4 #17 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo Clover text-to-video R@1 29.5 #17 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo Clover text-to-video R@10 66.3 #17 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo Clover text-to-video R@5 55.2 #17 of 26 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC Clover text-to-video Median Rank 24 #13 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC Clover text-to-video R@1 14.7 #13 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC Clover text-to-video R@10 38.2 #13 of 16 Archive leaderboard report
Zero-Shot Video Retrieval LSMDC Clover text-to-video R@5 29.2 #13 of 16 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT Clover text-to-video Median Rank 6 #25 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT Clover text-to-video R@1 26.4 #25 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT Clover text-to-video R@10 60 #25 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT Clover text-to-video R@5 49.5 #25 of 41 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

ALIGN

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections