Home › Datasets › task › Video Question Answering
Video Question Answering datasets
archive 2025-07-28
43 datasets carry the task tag "Video Question Answering" (the task itself: Video Question Answering), ordered by the archive's paper count. Page 1 of 1: 43 shown of 43. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Video Question Answering datasets 1–43 of 43
MSR-VTT (Microsoft Research Video to Text) is a large-scale dataset for the open domain video captioning, which consists of 10,000 video clips from 20 categories, and each video clip is annotated with 20 English sentences by Amazon…
640 papers · 8 benchmarks
HowTo100M is a large-scale dataset of narrated videos with an emphasis on instructional videos where content creators teach complex tasks with an explicit intention of explaining the visual content on screen.
286 papers · 1 benchmark
NExT-QA is a VideoQA benchmark targeting the explanation of video contents.
174 papers · 3 benchmarks
The ActivityNet-QA dataset contains 58,000 human-annotated QA pairs on 5,800 videos derived from the popular ActivityNet dataset.
146 papers · 2 benchmarks
The TVQA dataset is a large-scale video dataset for video question answering.
146 papers · 3 benchmarks
MVBench is a comprehensive Multi-modal Video understanding Benchmark.
139 papers · 3 benchmarks
The TGIF-QA dataset contains 165K QA pairs for the animated GIFs from the TGIF dataset [Li et al.
92 papers · 3 benchmarks
The MovieQA dataset is a dataset for movie question answering.
86 papers · 1 benchmark
The MSR-VTT-QA dataset is a benchmark for the task of Visual Question Answering (VQA) on the MSR-VTT (Microsoft Research Video to Text) dataset.
66 papers · 5 benchmarks
The MSVD-QA dataset is a Video Question Answering (VideoQA) dataset.
61 papers · 5 benchmarks
TVQA+ contains 310.8K bounding boxes, linking depicted objects to visual concepts in questions and answers.
60 papers · 0 benchmarks
The Tumblr GIF (TGIF) dataset contains 100K animated GIFs and 120K sentences describing visual content of the animated GIFs.
51 papers · 1 benchmark
AGQA (Action Genome Question Answering)
Action Genome Question Answering (AGQA) is a benchmark for compositional spatio-temporal reasoning.
30 papers · 0 benchmarks
VALUE (Video-And-Language Understanding Evaluation)
VALUE is a Video-And-Language Understanding Evaluation benchmark to test models that are generalizable to diverse tasks, domains, and datasets.
30 papers · 0 benchmarks
Video Instruction Dataset is used to train Video-ChatGPT.
30 papers · 7 benchmarks
To collect How2QA for video QA task, the same set of selected video clips are presented to another group of AMT workers for multichoice QA annotation.
28 papers · 2 benchmarks
We contribute an IntentQA dataset with diverse intents in daily social activities.
27 papers · 2 benchmarks
TVBench is a new benchmark specifically created to evaluate temporal understanding in video QA.
22 papers · 1 benchmark
iVQA (Instructional Video Question Answering)
An open-ended VideoQA benchmark that aims to: i) provide a well-defined evaluation by including five correct answer annotations per question and ii) avoid questions which can be answered without the video.
22 papers · 2 benchmarks
SUTD-TrafficQA (Singapore University of Technology and Design - Traffic Question Answering) is a dataset which takes the form of video QA based on 10,080 in-the-wild videos and annotated 62,535 QA pairs, for benchmarking the cognitive…
19 papers · 1 benchmark
EgoTask QA benchmark contains 40K balanced question-answer pairs selected from 368K programmatically generated questions generated over 2K egocentric videos.
17 papers · 1 benchmark
How to capture the present knowledge from surrounding situations and perform reasoning accordingly is crucial and challenging for machine intelligence.
17 papers · 2 benchmarks
The MSRVTT-MC (Multiple Choice) dataset is a video question-answering dataset created based on the MSR-VTT dataset.
14 papers · 2 benchmarks
OVBench is a benchmark tailored for real-time video understanding: - Memory, Perception, and Prediction of Temporal Contexts: Questions are framed to reference the present state of entities, requiring models to memorize/perceive/predict…
14 papers · 1 benchmark
The DramaQA focuses on two perspectives: 1) Hierarchical QAs as an evaluation metric based on the cognitive developmental stages of human intelligence.
12 papers · 1 benchmark
Neptune (Neptune Long Video Understanding Benchmark)
Neptune is a dataset consisting of challenging question-answer-decoy (QAD) sets for long videos (up to 15 minutes).
12 papers · 0 benchmarks
VLEP (Video-and-Language Event Prediction)
VLEP contains 28,726 future event prediction examples (along with their rationales) from 10,234 diverse TV Show and YouTube Lifestyle Vlog video clips.
11 papers · 1 benchmark
Perception Test is a benchmark designed to evaluate the perception and reasoning skills of multimodal models.
10 papers · 3 benchmarks
KnowIT VQA is a video dataset with 24,282 human-generated question-answer pairs about The Big Bang Theory.
9 papers · 0 benchmarks
VLM²-Bench: Benchmarking Vision-Language Models on Visual Cue Matching Description VLM²-Bench is the first comprehensive benchmark designed to evaluate vision-language models' (VLMs) ability to visually link matching cues across…
9 papers · 1 benchmark
A quantitative benchmark for developing and understanding video of fill-in-the-blank question-answering dataset with over 300,000 examples, based on descriptive video annotations for the visually impaired.
6 papers · 0 benchmarks
A dataset of 69,270,581 video clip, question and answer triplets (v, q, a).
5 papers · 0 benchmarks
TutorialVQA is a new type of dataset used to find answer spans in tutorial videos.
5 papers · 0 benchmarks
Text and signs around roads provide crucial information for drivers, vital for safe navigation and situational awareness.
4 papers · 1 benchmark
Video Localized Narratives is a new form of multimodal video annotations connecting vision and language.
3 papers · 0 benchmarks
CRIPP-VQA (Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering)
CRIPP-VQA is a video question answering dataset for reasoning about the implicit physical properties of objects in a scene.
2 papers · 0 benchmarks
CinePile is a question-answering-based, long-form video understanding dataset.
1 paper · 1 benchmark
Short-Films 20K (SF20K) is the largest publicly available movie dataset.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
VCG+112K (Video Instruction Dataset 112K)
Video-ChatGPT introduces the VideoInstruct100K dataset, which employs a semi-automatic annotation pipeline to generate 75K instruction-tuning QA pairs.
1 paper · 0 benchmarks
WildQA is a video understanding dataset of videos recorded in outside settings.
1 paper · 1 benchmark
Vript (🎬 Vript: A Video Is Worth Thousands of Words)
We construct a fine-grained video-text dataset with 12K annotated high-resolution videos (~400k clips).
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.