Papers › VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

28 Sep 2021EMNLP 2021 11arXiv:2109.14084archive 2025-07-28

Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, Christoph Feichtenhofer

We present VideoCLIP, a contrastive approach to pre-train a unified model for zero-shot video and text understanding, without using any labels on downstream tasks. VideoCLIP trains a transformer for video and text by contrasting temporally overlapping positive video-text pairs with hard negatives from nearest neighbor retrieval. Our experiments on a diverse series of downstream tasks, including sequence-level text-video retrieval, VideoQA, token-level action localization, and action segmentation reveal state-of-the-art performance, surpassing prior work, and in some cases even outperforming supervised approaches. Code is made available at https://github.com/pytorch/fairseq/tree/main/examples/MMPT.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

pytorch/fairseq officialmentioned in papermentioned on GitHubpytorch report
facebookresearch/fairseq mentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action LocalizationAction SegmentationLong Video Retrieval (Background Removed)RetrievalTemporal Action LocalizationTemporal Relation ExtractionVideo RetrievalZero-Shot Video Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Segmentation COIN VideoClip Frame accuracy 68.7 #4 of 9 Archive leaderboard report
Long Video Retrieval (Background Removed) YouCook2 VideoCLIP Cap. Avg. R@1 74.5 #3 of 6 Archive leaderboard report
Long Video Retrieval (Background Removed) YouCook2 VideoCLIP Cap. Avg. R@10 97.9 #3 of 6 Archive leaderboard report
Long Video Retrieval (Background Removed) YouCook2 VideoCLIP Cap. Avg. R@5 94.5 #3 of 6 Archive leaderboard report
Long Video Retrieval (Background Removed) YouCook2 VideoCLIP DTW R@1 56.0 #3 of 6 Archive leaderboard report
Long Video Retrieval (Background Removed) YouCook2 VideoCLIP DTW R@10 89.9 #3 of 6 Archive leaderboard report
Long Video Retrieval (Background Removed) YouCook2 VideoCLIP DTW R@5 96.3 #3 of 6 Archive leaderboard report
Long Video Retrieval (Background Removed) YouCook2 VideoCLIP OTAM R@1 52.8 #3 of 6 Archive leaderboard report
Long Video Retrieval (Background Removed) YouCook2 VideoCLIP OTAM R@10 89.2 #3 of 6 Archive leaderboard report
Long Video Retrieval (Background Removed) YouCook2 VideoCLIP OTAM R@5 95.0 #3 of 6 Archive leaderboard report
Temporal Action Localization CrossTask VideoCLIP Recall 47.3 #1 of 7 Archive leaderboard report
Temporal Relation Extraction Vinoground VideoCLIP Group Score 1.2 #22 of 24 Archive leaderboard report
Temporal Relation Extraction Vinoground VideoCLIP Text Score 17 #22 of 24 Archive leaderboard report
Temporal Relation Extraction Vinoground VideoCLIP Video Score 2.8 #22 of 24 Archive leaderboard report
Video Retrieval MSR-VTT-1kA VideoCLIP text-to-video R@1 30.9 #50 of 63 Archive leaderboard report
Video Retrieval MSR-VTT-1kA VideoCLIP text-to-video R@10 66.8 #50 of 63 Archive leaderboard report
Video Retrieval MSR-VTT-1kA VideoCLIP text-to-video R@5 55.4 #50 of 63 Archive leaderboard report
Video Retrieval YouCook2 VideoCLIP text-to-video R@1 32.2 #3 of 16 Archive leaderboard report
Video Retrieval YouCook2 VideoCLIP text-to-video R@10 75.0 #3 of 16 Archive leaderboard report
Video Retrieval YouCook2 VideoCLIP text-to-video R@5 62.6 #3 of 16 Archive leaderboard report
Video Retrieval YouCook2 VideoCLIP (zero-shot) text-to-video R@1 22.7 #8 of 16 Archive leaderboard report
Video Retrieval YouCook2 VideoCLIP (zero-shot) text-to-video R@10 63.1 #8 of 16 Archive leaderboard report
Video Retrieval YouCook2 VideoCLIP (zero-shot) text-to-video R@5 50.4 #8 of 16 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo VideoCLIP text-to-video R@1 16.6 #26 of 26 Archive leaderboard report
Zero-Shot Video Retrieval DiDeMo VideoCLIP text-to-video R@5 46.9 #26 of 26 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT VideoCLIP text-to-video R@1 10.4 #36 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT VideoCLIP text-to-video R@10 30.0 #36 of 41 Archive leaderboard report
Zero-Shot Video Retrieval MSR-VTT VideoCLIP text-to-video R@5 22.2 #36 of 41 Archive leaderboard report
Zero-Shot Video Retrieval YouCook2 VideoCLIP text-to-video R@1 22.7 #3 of 9 Archive leaderboard report
Zero-Shot Video Retrieval YouCook2 VideoCLIP text-to-video R@10 63.1 #3 of 9 Archive leaderboard report
Zero-Shot Video Retrieval YouCook2 VideoCLIP text-to-video R@5 50.4 #3 of 9 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections