Home › Datasets › task › Zero-Shot Video Question Answer
Zero-Shot Video Question Answer datasets
archive 2025-07-28
18 datasets carry the task tag "Zero-Shot Video Question Answer" (the task itself: Zero-Shot Video Question Answer), ordered by the archive's paper count. Page 1 of 1: 18 shown of 18. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Zero-Shot Video Question Answer datasets 1–18 of 18
NExT-QA is a VideoQA benchmark targeting the explanation of video contents.
174 papers · 3 benchmarks
Video-MME stands for Video Multi-Modal Evaluation.
152 papers · 2 benchmarks
The ActivityNet-QA dataset contains 58,000 human-annotated QA pairs on 5,800 videos derived from the popular ActivityNet dataset.
146 papers · 2 benchmarks
The TVQA dataset is a large-scale video dataset for video question answering.
146 papers · 3 benchmarks
MVBench is a comprehensive Multi-modal Video understanding Benchmark.
139 papers · 3 benchmarks
EgoSchema is very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems.
112 papers · 3 benchmarks
The TGIF-QA dataset contains 165K QA pairs for the animated GIFs from the TGIF dataset [Li et al.
92 papers · 3 benchmarks
The MSR-VTT-QA dataset is a benchmark for the task of Visual Question Answering (VQA) on the MSR-VTT (Microsoft Research Video to Text) dataset.
66 papers · 5 benchmarks
The MSVD-QA dataset is a Video Question Answering (VideoQA) dataset.
61 papers · 5 benchmarks
TVQA+ contains 310.8K bounding boxes, linking depicted objects to visual concepts in questions and answers.
60 papers · 0 benchmarks
We contribute an IntentQA dataset with diverse intents in daily social activities.
27 papers · 2 benchmarks
We study visually grounded VideoQA in response to the emerging trends of utilizing pretraining techniques for video-language understanding.
23 papers · 1 benchmark
How to capture the present knowledge from surrounding situations and perform reasoning accordingly is crucial and challenging for machine intelligence.
17 papers · 2 benchmarks
Neptune (Neptune Long Video Understanding Benchmark)
Neptune is a dataset consisting of challenging question-answer-decoy (QAD) sets for long videos (up to 15 minutes).
12 papers · 0 benchmarks
VNBench is a comprehensive benchmark suite for video generative models, which evaluates video generation quality across specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods.
11 papers · 1 benchmark
Large multimodal models (LMMs) are processing increasingly longer and richer inputs.
5 papers · 1 benchmark
A short clip of video may contain progression of multiple events and an interesting story line.
3 papers · 3 benchmarks
CinePile is a question-answering-based, long-form video understanding dataset.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.