Home › Datasets › modality › Videos
Videos datasets
archive 2025-07-28
1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 5 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Videos datasets 193–240 of 1,014
Contains 68,536 activity instances in 68.8 hours of first and third-person video, making it one of the largest and most diverse egocentric datasets available.
33 papers · 1 benchmark
A new multitask action quality assessment (AQA) dataset, the largest to date, comprising of more than 1600 diving samples; contains detailed annotations for fine-grained action recognition, commentary generation, and estimating the AQA…
33 papers · 2 benchmarks
WMCA (Wide Multi Channel Presentation Attack)
The Wide Multi Channel Presentation Attack (WMCA) database consists of 1941 short video recordings of both bonafide and presentation attacks from 72 different identities.
33 papers · 1 benchmark
AUTSL (Ankara University Turkish Sign Language Dataset)
The Ankara University Turkish Sign Language Dataset (AUTSL) is a large-scale, multimode dataset that contains isolated Turkish sign videos.
32 papers · 1 benchmark
CholecT50 is a dataset of endoscopic videos of laparoscopic cholecystectomy surgery introduced to enable research on fine-grained action recognition in laparoscopic surgery.
32 papers · 5 benchmarks
EMDB contains in-the-wild videos of human activity recorded with a hand-held iPhone.
32 papers · 2 benchmarks
Ego4D is a massive-scale egocentric video dataset and benchmark suite.
32 papers · 6 benchmarks
SportsMOT (SportsMOT: A Large Multi-Object Tracking Dataset in Multiple Sports Scenes)
Motivation Multi-object tracking (MOT) is a fundamental task in computer vision, aiming to estimate objects (e.g., pedestrians and vehicles) bounding boxes and identities in video sequences.
32 papers · 3 benchmarks
V2X-Sim, short for vehicle-to-everything simulation, is the a synthetic collaborative perception dataset in autonomous driving developed by AI4CE Lab at NYU and MediaBrain Group at SJTU to facilitate collaborative perception between…
32 papers · 1 benchmark
MAFW is a large-scale, multi-modal, compound affective database for dynamic facial expression recognition in the wild.
31 papers · 2 benchmarks
A large scale dataset with daily-living activities performed in a natural manner.
31 papers · 1 benchmark
AGQA (Action Genome Question Answering)
Action Genome Question Answering (AGQA) is a benchmark for compositional spatio-temporal reasoning.
30 papers · 0 benchmarks
VALUE (Video-And-Language Understanding Evaluation)
VALUE is a Video-And-Language Understanding Evaluation benchmark to test models that are generalizable to diverse tasks, domains, and datasets.
30 papers · 0 benchmarks
VGG-SS (VGG Sound Source) is a benchmark for evaluating sound source localisation in videos.
30 papers · 0 benchmarks
Video Instruction Dataset is used to train Video-ChatGPT.
30 papers · 7 benchmarks
The NVGesture dataset focuses on touchless driver controlling.
29 papers · 1 benchmark
Spring (Spring: A High-Resolution High-Detail Dataset and Benchmark for Scene Flow, Optical Flow and Stereo)
Spring is a large, high-resolution and high-detail, computer-generated benchmark for scene flow, optical flow, and stereo.
29 papers · 3 benchmarks
A realistic dataset composed of 27 episodes from 6 popular TV series.
29 papers · 1 benchmark
Dynamic FAUST extends the FAUST dataset to dynamic 4D data.
28 papers · 1 benchmark
The Gaming 3D Dataset (G3D) focuses on real-time action recognition in a gaming scenario.
28 papers · 2 benchmarks
To collect How2QA for video QA task, the same set of selected video clips are presented to another group of AMT workers for multichoice QA annotation.
28 papers · 2 benchmarks
KITTI MOTS (KITTI Multi-Object Tracking and Segmentation (MOTS) Evaluation)
The Multi-Object and Segmentation (MOTS) benchmark [2] consists of 21 training sequences and 29 test sequences.
28 papers · 1 benchmark
The MannequinChallenge Dataset (MQC) provides in-the-wild videos of people in static poses while a hand-held camera pans around the scene.
28 papers · 0 benchmarks
The Oxford Radar RobotCar Dataset is a radar extension to The Oxford RobotCar Dataset.
28 papers · 4 benchmarks
DeeperForensics-1.0 represents the largest face forgery detection dataset by far, with 60,000 videos constituted by a total of 17.6 million frames, 10 times larger than existing datasets of the same kind.
27 papers · 0 benchmarks
We contribute an IntentQA dataset with diverse intents in daily social activities.
27 papers · 2 benchmarks
KoDF (Korean DeepFake Detection Dataset)
The Korean DeepFake Detection Dataset (KoDF) is a large-scale collection of synthesized and real videos focused on Korean subjects, used for the task of deepfake detection.
27 papers · 0 benchmarks
RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song)
The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) contains 7,356 files (total size: 24.8 GB).
27 papers · 6 benchmarks
ROAD (ROAD: The ROad event Awareness Dataset for Autonomous Driving)
ROAD is designed to test an autonomous vehicle's ability to detect road events, defined as triplets composed by an active agent, the action(s) it performs and the corresponding scene locations.
27 papers · 0 benchmarks
UVO (Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation)
UVO is a new benchmark for open-world class-agnostic object segmentation in videos.
27 papers · 2 benchmarks
The Vid4 dataset is generally used for testing video super-resolution.
27 papers · 3 benchmarks
Dataset of high-resolution (4096×2160), high-fps (1000fps) video frames with extreme motion.
27 papers · 1 benchmark
Animal Kingdom is a large and diverse dataset that provides multiple annotated tasks to enable a more thorough understanding of natural animal behaviors.
26 papers · 2 benchmarks
CRVD (Captured Raw Video Denoising)
The CRVD dataset consists of 55 groups of noisy-clean videos with ISO values ranging from 1600 to 25600.
26 papers · 1 benchmark
Cityscapes-VPS is a video extension of the Cityscapes validation split.
26 papers · 1 benchmark
The Drive&Act dataset is a state of the art multi modal benchmark for driver behavior recognition.
26 papers · 1 benchmark
Our dataset was made of videos from MSU Video Upscalers Benchmark Dataset, MSU Video Super-Resolution Benchmark Dataset and MSU Super-Resolution for Video Compression Benchmark Dataset.
26 papers · 1 benchmark
Street Scene is a dataset for video anomaly detection.
26 papers · 3 benchmarks
BDD-A (Berkeley DeepDrive Attention)
Dataset Statistics: The statistics of our dataset are summarized and compared with the largest existing dataset (DR(eye)VE) [1] in Table 1.
25 papers · 0 benchmarks
A three million frame, multi-view, furniture assembly video dataset that includes depth, atomic actions, object segmentation, and human pose.
25 papers · 1 benchmark
This is a dataset for a video super-resolution task.
25 papers · 1 benchmark
The signing is recorded by a stationary color camera placed in front of the sign language interpreters.
25 papers · 1 benchmark
DHF1K is a video saliency dataset which contains a ground-truth map of binary pixel-wise gaze fixation points and a continuous map of the fixation points after being blurred by a gaussian filter.
24 papers · 1 benchmark
FineAction contains 103K temporal instances of 106 action categories, annotated in 17K untrimmed videos.
24 papers · 3 benchmarks
4D-DRESS (A 4D Dataset of Real-world Human Clothing with Semantic Annotations)
4D-DRESS is the first real-world 4D dataset of human clothing, capturing 64 human outfits in more than 520 motion sequences.
23 papers · 4 benchmarks
A large-scale 4D egocentric dataset with rich annotations, to catalyze the research of category-level human-object interaction.
23 papers · 0 benchmarks
MMAct is a large-scale dataset for multi/cross modal action understanding.
23 papers · 1 benchmark
We study visually grounded VideoQA in response to the emerging trends of utilizing pretraining techniques for video-language understanding.
23 papers · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.