Home › Datasets › task › Video Understanding

Video Understanding datasets

archive 2025-07-28

56 datasets carry the task tag "Video Understanding" (the task itself: Video Understanding), ordered by the archive's paper count. Page 1 of 2: 48 shown of 56. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Video Understanding datasets 1–48 of 56

Kinetics (Kinetics Human Action Video Dataset)
The Kinetics dataset is a large-scale, high-quality dataset for human action recognition in videos.
1,341 papers · 18 benchmarks
The Charades dataset is composed of 9,848 videos of daily indoors activities with an average length of 30 seconds, involving interactions with 46 objects classes in 15 types of indoor scenes and containing a vocabulary of 30 verbs leading…
428 papers · 6 benchmarks
Charades-STA is a new dataset built on top of Charades by adding sentence temporal annotations.
236 papers · 4 benchmarks
The ShanghaiTech Campus dataset has 13 scenes with complex light conditions and camera angles.
207 papers · 4 benchmarks
SEED-Bench consists of 19K multiple choice questions with accurate human annotations (~6 larger than existing benchmarks), which spans 12 evaluation dimensions including the comprehension of both the image and video modality.
137 papers · 0 benchmarks
AVA (Atomic Visual Actions)
AVA is a project that provides audiovisual annotations of video for improving our understanding of human activity.
113 papers · 7 benchmarks
A novel large-scale corpus of manual annotations for the SoccerNet video dataset, along with open challenges to encourage more research in soccer understanding and broadcast production.
58 papers · 6 benchmarks
MovieNet is a holistic dataset for movie understanding.
54 papers · 1 benchmark
InternVid is a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodAL understanding and generation.
46 papers · 0 benchmarks
The EPIC-KITCHENS-55 dataset comprises a set of 432 egocentric videos recorded by 32 participants in their kitchens at 60fps with a head mounted camera.
42 papers · 3 benchmarks
Contains 68,536 activity instances in 68.8 hours of first and third-person video, making it one of the largest and most diverse egocentric datasets available.
33 papers · 1 benchmark
A new multitask action quality assessment (AQA) dataset, the largest to date, comprising of more than 1600 diving samples; contains detailed annotations for fine-grained action recognition, commentary generation, and estimating the AQA…
33 papers · 2 benchmarks
CCD (Car Crash Dataset)
Car Crash Dataset (CCD) is collected for traffic accident analysis.
18 papers · 1 benchmark
VidSitu is a dataset for the task of semantic role labeling in videos (VidSRL).
18 papers · 0 benchmarks
STAR Benchmark (Situated Reasoning)
How to capture the present knowledge from surrounding situations and perform reasoning accordingly is crucial and challenging for machine intelligence.
17 papers · 2 benchmarks
HVU (Holistic Video Understanding)
HVU is organized hierarchically in a semantic taxonomy that focuses on multi-label and multi-task video understanding as a comprehensive problem that encompasses the recognition of multiple semantic aspects in the dynamic scene.
16 papers · 0 benchmarks
A large-scale dataset for retrieval and event localisation in video.
15 papers · 1 benchmark
OVBench is a benchmark tailored for real-time video understanding: - Memory, Perception, and Prediction of Temporal Contexts: Questions are framed to reference the present state of entities, requiring models to memorize/perceive/predict…
14 papers · 1 benchmark
Aesthetic Visual Analysis is a dataset for aesthetic image assessment that contains over 250,000 images along with a rich variety of meta-data including a large number of aesthetic scores for each image, semantic labels for over 60…
12 papers · 1 benchmark
The DramaQA focuses on two perspectives: 1) Hierarchical QAs as an evaluation metric based on the cognitive developmental stages of human intelligence.
12 papers · 1 benchmark
ImageCoDe (Image Retrieval from Contextual Descriptions)
Given 10 minimally contrastive (highly similar) images and a complex description for one of them, the task is to retrieve the correct image.
11 papers · 1 benchmark
Moviescope is a large-scale dataset of 5,000 movies with corresponding video trailers, posters, plots and metadata.
6 papers · 0 benchmarks
V2C (Video-to-Commonsense)
6 papers · 0 benchmarks
ChronoMagic with 2265 metamorphic time-lapse videos, each accompanied by a detailed caption.
5 papers · 0 benchmarks
InfiniBench (InfiniBench: A Comprehensive Benchmark for Large Multimodal Models in Very Long Video Understanding)
We introduce InfiniBench a comprehensive benchmark for very long video understanding, which presents 1) The longest video duration, averaging 76.34 minutes; 2) The largest number of question-answer pairs, 108.2K; 3) Diversity in questions…
5 papers · 0 benchmarks
The MLB-YouTube dataset is a new, large-scale dataset consisting of 20 baseball games from the 2017 MLB post-season available on YouTube with over 42 hours of video footage.
5 papers · 0 benchmarks
This dataset has the following citation: M.
5 papers · 1 benchmark
Comprises of 171,191 video segments from 346 high-quality soccer games.
5 papers · 0 benchmarks
M³-VOS (M³-VOS: Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation)
💡 Description A new benchmark, Multi-Phase, Multi-Transition, and Multi-Scenery Video Object Segmentation (M³-VOS), to verify the ability of models to understand object phases, which consists of 479 high-resolution videos spanning over 10…
4 papers · 1 benchmark
A novel dataset that represents complex conversational interactions between two individuals via 3D pose.
3 papers · 0 benchmarks
DeepSportradar is a benchmark suite of computer vision tasks, datasets and benchmarks for automated sport understanding.
3 papers · 0 benchmarks
Fitness-AQA (Fitness Action Quality Assessment [ECCV 2022])
Largest, first-of-its-kind, in-the-wild, fine-grained workout/exercise posture analysis dataset, covering three different exercises: BackSquat, Barbell Row, and Overhead Press.
3 papers · 0 benchmarks
A multilingual, multimodal and multi-aspect, expertly-annotated dataset of diverse short videos extracted from short-video social media platform - Moj.
2 papers · 0 benchmarks
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
2 papers · 0 benchmarks
Cinescale (CineScale: A dataset of cinematic shot scale in movies)
We provide a database containing shot scale annotations (i.e., the apparent distance of the camera from the subject of a filmed scene) for more than 792,000 image frames.
2 papers · 0 benchmarks
Contains annotations of human activity with different sub-actions, e.g., activity Ping-Pong with four sub-actions which are pickup-ball, hit, bounce-ball and serve.
2 papers · 0 benchmarks
SVBench (Streaming Video Understanding Benchmark)
Dataset Card for SVBench This dataset card aims to provide a comprehensive overview of the SVBench dataset, including its purpose, structure, and sources.
2 papers · 0 benchmarks
Stanford-ECM is an egocentric multimodal dataset which comprises about 27 hours of egocentric video augmented with heart rate and acceleration data.
2 papers · 0 benchmarks
VTC (Videos, Titles and Comments)
VTC is a large-scale multimodal dataset containing video-caption pairs (~300k) alongside comments that can be used for multimodal representation learning.
2 papers · 0 benchmarks
This large collection of over 161,000 video-label pairs of video clips, shows humans drawing letters and digits in the air, and is used to evaluate a model’s ability to classify articulated motions correctly.
1 paper · 0 benchmarks
The feature files are named with the youtube IDs.
1 paper · 0 benchmarks
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
1 paper · 0 benchmarks
This spatio-temporal actions dataset for video understanding consists of 4 parts: original videos, cropped videos, video frames, and annotation files.
1 paper · 0 benchmarks
Kinetics-GEB+ (Generic Event Boundary Captioning, Grounding and Retrieval) is a dataset that consists of over 170k boundaries associated with captions describing status changes in the generic events in 12K videos.
1 paper · 3 benchmarks
LSDBench (Long-video Sampling Dilemma Benchmark)
A benchmark that focuses on the sampling dilemma in long-video tasks.
1 paper · 0 benchmarks
MuseChat Dataset (MuseChat: A Conversational Music Recommendation System for Videos (CVPR 2024 Highlight Paper))
Music recommendation for videos attracts growing interest in multi-modal research.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.