Home › Datasets › task › Video Captioning
Video Captioning datasets
archive 2025-07-28
38 datasets carry the task tag "Video Captioning" (the task itself: Video Captioning), ordered by the archive's paper count. Page 1 of 1: 38 shown of 38. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Video Captioning datasets 1–38 of 38
Kinetics (Kinetics Human Action Video Dataset)
The Kinetics dataset is a large-scale, high-quality dataset for human action recognition in videos.
1,341 papers · 18 benchmarks
MSR-VTT (Microsoft Research Video to Text) is a large-scale dataset for the open domain video captioning, which consists of 10,000 video clips from 20 categories, and each video clip is annotated with 20 English sentences by Amazon…
640 papers · 8 benchmarks
MSVD (Microsoft Research Video Description Corpus)
The Microsoft Research Video Description Corpus (MSVD) dataset consists of about 120K sentences collected during the summer of 2010.
327 papers · 3 benchmarks
HowTo100M is a large-scale dataset of narrated videos with an emphasis on instructional videos where content creators teach complex tasks with an explicit intention of explaining the visual content on screen.
286 papers · 1 benchmark
WebVid contains 10 million video clips with captions, sourced from the web.
257 papers · 1 benchmark
The ActivityNet Captions dataset is built on ActivityNet v1.3 which includes 20k YouTube untrimmed videos with 100k caption annotations.
255 papers · 6 benchmarks
YouCook2 is the largest task-oriented, instructional video dataset in the vision community.
198 papers · 7 benchmarks
VATEX is multilingual, large, linguistically complex, and diverse dataset in terms of both video and natural language descriptions.
118 papers · 3 benchmarks
The Tumblr GIF (TGIF) dataset contains 100K animated GIFs and 120K sentences describing visual content of the animated GIFs.
51 papers · 1 benchmark
This data set was prepared from 88 open-source YouTube cooking videos.
45 papers · 0 benchmarks
A new multitask action quality assessment (AQA) dataset, the largest to date, comprising of more than 1600 diving samples; contains detailed annotations for fine-grained action recognition, commentary generation, and estimating the AQA…
33 papers · 2 benchmarks
VALUE (Video-And-Language Understanding Evaluation)
VALUE is a Video-And-Language Understanding Evaluation benchmark to test models that are generalizable to diverse tasks, domains, and datasets.
30 papers · 0 benchmarks
To collect How2QA for video QA task, the same set of selected video clips are presented to another group of AMT workers for multichoice QA annotation.
28 papers · 2 benchmarks
Violin (VIdeO-and-Language INference)
Video-and-Language Inference is the task of joint multimodal understanding of video and text.
18 papers · 0 benchmarks
TV show Caption is a large-scale multimodal captioning dataset, containing 261,490 caption descriptions paired with 108,965 short video moments.
16 papers · 1 benchmark
MaSS (Multilingual corpus of Sentence-aligned Spoken utterances) is an extension of the CMU Wilderness Multilingual Speech Dataset, a speech dataset based on recorded readings of the New Testament.
15 papers · 3 benchmarks
ViTT (Video Timeline Tags)
The ViTT dataset consists of human produced segment-level annotations for 8,169 videos.
14 papers · 2 benchmarks
EgoExoLearn is a fascinating dataset designed to bridge the gap between egocentric and exocentric views of procedural activities.
12 papers · 3 benchmarks
The AI City Challenge, hosted at CVPR 2024, focuses on harnessing AI to enhance operational efficiency in physical settings such as retail and warehouse environments, and Intelligent Traffic Systems (ITS).
10 papers · 1 benchmark
Amazon Mechanical Turk (AMT) is used to collect annotations on HowTo100M videos.
9 papers · 0 benchmarks
V2C (Video-to-Commonsense)
6 papers · 0 benchmarks
SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model
5 papers · 0 benchmarks
VidChapters-7M is a dataset of 817K user-chaptered videos including 7M chapters in total.
5 papers · 4 benchmarks
MSRVTT-CTN Dataset This dataset contains CTN annotations for the MSRVTT-CTN benchmark dataset in JSON format.
4 papers · 1 benchmark
MSVD-CTN (MSVD Causal-Temporal Narrative)
MSVD-CTN Dataset This dataset contains CTN annotations for the MSVD-CTN benchmark dataset in JSON format.
4 papers · 1 benchmark
Youku-mPLUG is a large Chinese high-quality video-language dataset which is collected from Youku.com, a well-known Chinese video-sharing website, with strict criteria of safety, diversity, and quality.
4 papers · 0 benchmarks
The dataset contains the annotations of characters' visual appearances, in the form of tracks of face bounding boxes, and the associations with characters' textual mentions, when available.
3 papers · 1 benchmark
A short clip of video may contain progression of multiple events and an interesting story line.
3 papers · 3 benchmarks
This dataset is the Hindi version of standard English MSR-VTT dataset.
2 papers · 1 benchmark
Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks.
2 papers · 0 benchmarks
We provide a dataset called MMAC Captions for sensor-augmented egocentric-video captioning.
2 papers · 0 benchmarks
NSVA (NBA dataset for Sports Video Analysis)
NVSA is a large-scale NBA dataset for Sports Video Analysis (NSVA) with a focus on sports video captioning.
2 papers · 0 benchmarks
ChinaOpen is a new video dataset targeted at open-world multimodal learning, with raw data gathered from Bilibili, a popular Chinese video-sharing website.
1 paper · 1 benchmark
Kinetics-GEB+ (Generic Event Boundary Captioning, Grounding and Retrieval) is a dataset that consists of over 170k boundaries associated with captions describing status changes in the generic events in 12K videos.
1 paper · 3 benchmarks
The Large Scale Movie Description Challenge (LSMDC) - Context is an augmented version of the original LSMDC dataset with movie scripts as contextual text.
1 paper · 0 benchmarks
MSVD-Indonesian is derived from the MSVD dataset, which is obtained with the help of a machine translation service.
1 paper · 3 benchmarks
W-Oops consists of 2,100 unintentional human action videos, with 44 goal-directed and 30 unintentional video-level activity labels collected through human annotations.
1 paper · 0 benchmarks
Vript (🎬 Vript: A Video Is Worth Thousands of Words)
We construct a fine-grained video-text dataset with 12K annotated high-resolution videos (~400k clips).
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.