Home › Datasets › task › Video Retrieval

Video Retrieval datasets

archive 2025-07-28

35 datasets carry the task tag "Video Retrieval" (the task itself: Video Retrieval), ordered by the archive's paper count. Page 1 of 1: 35 shown of 35. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Video Retrieval datasets 1–35 of 35

Kinetics (Kinetics Human Action Video Dataset)
The Kinetics dataset is a large-scale, high-quality dataset for human action recognition in videos.
1,341 papers · 18 benchmarks
The ActivityNet dataset contains 200 different types of activities and a total of 849 hours of videos collected from YouTube.
807 papers · 17 benchmarks
MSR-VTT (Microsoft Research Video to Text) is a large-scale dataset for the open domain video captioning, which consists of 10,000 video clips from 20 categories, and each video clip is annotated with 20 English sentences by Amazon…
640 papers · 8 benchmarks
MSVD (Microsoft Research Video Description Corpus)
The Microsoft Research Video Description Corpus (MSVD) dataset consists of about 120K sentences collected during the summer of 2010.
327 papers · 3 benchmarks
HowTo100M is a large-scale dataset of narrated videos with an emphasis on instructional videos where content creators teach complex tasks with an explicit intention of explaining the visual content on screen.
286 papers · 1 benchmark
WebVid contains 10 million video clips with captions, sourced from the web.
257 papers · 1 benchmark
Charades-STA is a new dataset built on top of Charades by adding sentence temporal annotations.
236 papers · 4 benchmarks
DiDeMo (Distinct Describable Moments)
The Distinct Describable Moments (DiDeMo) dataset is one of the largest and most diverse datasets for the temporal localization of events in videos given natural language descriptions.
216 papers · 3 benchmarks
YouCook2 is the largest task-oriented, instructional video dataset in the vision community.
198 papers · 7 benchmarks
LSMDC (Large Scale Movie Description Challenge)
This dataset contains 118,081 short video clips extracted from 202 movies.
126 papers · 3 benchmarks
VATEX (Video And TEXt)
VATEX is multilingual, large, linguistically complex, and diverse dataset in terms of both video and natural language descriptions.
118 papers · 3 benchmarks
TGIF (Tumblr GIF)
The Tumblr GIF (TGIF) dataset contains 100K animated GIFs and 120K sentences describing visual content of the animated GIFs.
51 papers · 1 benchmark
TVR (TV show Retrieval)
A new multimodal retrieval dataset.
34 papers · 2 benchmarks
To collect How2QA for video QA task, the same set of selected video clips are presented to another group of AMT workers for multichoice QA annotation.
28 papers · 2 benchmarks
CMD (Condensed Movies Dataset)
Consists of the key scenes from over 3K movies: each key scene is accompanied by a high level semantic description of the scene, character face-tracks, and metadata about the movie.
19 papers · 0 benchmarks
Violin (VIdeO-and-Language INference)
Video-and-Language Inference is the task of joint multimodal understanding of video and text.
18 papers · 0 benchmarks
The FIVR-200K dataset has been collected to simulate the problem of Fine-grained Incident Video Retrieval (FIVR).
16 papers · 1 benchmark
TVC (TV show Captions)
TV show Caption is a large-scale multimodal captioning dataset, containing 261,490 caption descriptions paired with 108,965 short video moments.
16 papers · 1 benchmark
A large-scale video dataset, featuring clips from movies with detailed captions.
15 papers · 1 benchmark
A large-scale dataset for retrieval and event localisation in video.
15 papers · 1 benchmark
EgoExoLearn is a fascinating dataset designed to bridge the gap between egocentric and exocentric views of procedural activities.
12 papers · 3 benchmarks
Amazon Mechanical Turk (AMT) is used to collect annotations on HowTo100M videos.
9 papers · 0 benchmarks
MVK (Marine Video Kit)
The dataset contains single-shot videos taken from moving cameras in underwater environments.
9 papers · 1 benchmark
HiREST (HIerarchical REtrieval and STep-captioning)
HiREST (HIerarchical REtrieval and STep-captioning) dataset is a benchmark that covers hierarchical information retrieval and visual/textual stepwise summarization from an instructional video corpus.
7 papers · 0 benchmarks
ChronoMagic with 2265 metamorphic time-lapse videos, each accompanied by a detailed caption.
5 papers · 0 benchmarks
A short clip of video may contain progression of multiple events and an interesting story line.
3 papers · 3 benchmarks
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
2 papers · 0 benchmarks
VTC (Videos, Titles and Comments)
VTC is a large-scale multimodal dataset containing video-caption pairs (~300k) alongside comments that can be used for multimodal representation learning.
2 papers · 0 benchmarks
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
1 paper · 0 benchmarks
IAW Dataset (Ikea Assembly In The Wild Dataset)
The IAW dataset contains 420 Ikea furniture pieces from 14 common categories e.g.
1 paper · 0 benchmarks
Kinetics-GEB+ (Generic Event Boundary Captioning, Grounding and Retrieval) is a dataset that consists of over 170k boundaries associated with captions describing status changes in the generic events in 12K videos.
1 paper · 3 benchmarks
MSVD-Indonesian is derived from the MSVD dataset, which is obtained with the help of a machine translation service.
1 paper · 3 benchmarks
STVD-PVCD (Partial Video Copy Detection Dataset)
STVD is the largest public dataset on the PVCD task.
1 paper · 1 benchmark
VILT (Video Instructions Linking for Complex Tasks)
VILT is a new benchmark collection of tasks and multimodal video content.
1 paper · 0 benchmarks
First of its kind paired win-fail action understanding dataset with samples from the following domains: “General Stunts,” “Internet Wins-Fails,” “Trick Shots,” & “Party Games.” The task is to identify successful and failed attempts at…
1 paper · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.