Home › Datasets › modality › Videos
Videos datasets
archive 2025-07-28
1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 8 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Videos datasets 337–384 of 1,014
RareAct is a video dataset of unusual actions, including actions like “blend phone”, “cut keyboard” and “microwave shoes”.
12 papers · 1 benchmark
The Robotic Pushing Dataset is a dataset for video prediction for real-world interactive agents which consists of 59,000 robot interactions involving pushing motions, including a test set with novel objects.
12 papers · 0 benchmarks
Co-speech gestures are everywhere.
12 papers · 1 benchmark
The UCFRep dataset contains 526 annotated repetitive action videos.
12 papers · 1 benchmark
VOT2014 (Visual Object Tracking Challenge 2014)
The dataset comprises 25 short sequences showing various objects in challenging backgrounds.
12 papers · 1 benchmark
VideoMatte240K consists of 484 high-resolution green screen videos and generate a total of 240,709 unique frames of alpha mattes and foregrounds with chroma-key software Adobe After Effects.
12 papers · 0 benchmarks
VideoSet is a large-scale compressed video quality dataset based on just-noticeable-difference (JND) measurement.
12 papers · 0 benchmarks
YT-UGC is a large scale UGC (User Generated Content) dataset (1,500 20 sec video clips) sampled from millions of YouTube videos.
12 papers · 0 benchmarks
4D-OR includes a total of 6734 scenes, recorded by six calibrated RGB-D Kinect sensors 1 mounted to the ceiling of the OR, with one frame-per-second, providing synchronized RGB and depth images.
11 papers · 3 benchmarks
BL30K is a synthetic dataset rendered using Blender with ShapeNet's data.
11 papers · 0 benchmarks
This is the home of a collaborative data collection effort by U.
11 papers · 1 benchmark
This is the home of a collaborative data collection effort by U.
11 papers · 1 benchmark
the YF-E6 emotion dataset using the 6 basic emotion type as keywords on social video-sharing websites including YouTube and Flickr, leading to a total of 3000 videos.
11 papers · 1 benchmark
GSL (Greek Sign Language)
Dataset Description The Greek Sign Language (GSL) is a large-scale RGB+D dataset, suitable for Sign Language Recognition (SLR) and Sign Language Translation (SLT).
11 papers · 1 benchmark
The GTA Indoor Motion dataset (GTA-IM) that emphasizes human-scene interactions in the indoor environments.
11 papers · 2 benchmarks
The human-Related version of the CUHK Avenue dataset, first presented by Morais et al.
11 papers · 1 benchmark
OpenLane-V2 is the world's first perception and reasoning benchmark for scene structure in autonomous driving.
11 papers · 2 benchmarks
PATS (Pose Audio Transcript Style)
PATS dataset consists of a diverse and large amount of aligned pose, audio and transcripts.
11 papers · 0 benchmarks
SLOPER4D is a novel scene-aware dataset collected in large urban environments to facilitate the research of global human pose estimation (GHPE) with human-scene interaction in the wild.
11 papers · 1 benchmark
TRIPOD (TuRnIng POint Dataset)
TRIPOD contains screenplays and plot synopses with turning point (TP) annotations for 99 movies.
11 papers · 0 benchmarks
VLEP (Video-and-Language Event Prediction)
VLEP contains 28,726 future event prediction examples (along with their rationales) from 10,234 diverse TV Show and YouTube Lifestyle Vlog video clips.
11 papers · 1 benchmark
VOST consists of more than 700 high-resolution videos, captured in diverse environments, which are 20 seconds long on average and densely labeled with instance masks.
11 papers · 0 benchmarks
The Video2GIF dataset contains over 100,000 pairs of GIFs and their source videos.
11 papers · 0 benchmarks
iMiGUE is a dataset for emotional artificial intelligence research: identity-free video dataset for Micro-Gesture Understanding and Emotion analysis (iMiGUE).
11 papers · 1 benchmark
The AI City Challenge, hosted at CVPR 2024, focuses on harnessing AI to enhance operational efficiency in physical settings such as retail and warehouse environments, and Intelligent Traffic Systems (ITS).
10 papers · 1 benchmark
CalMS21 (Caltech Mouse Social Interactions)
The Caltech Mouse Social Interactions (CalMS21) dataset is a multi-agent dataset from behavioral neuroscience.
10 papers · 0 benchmarks
DDPM (Deception Detection and Physiological Monitoring)
The Deception Detection and Physiological Monitoring (DDPM) dataset captures an interview scenario in which the interviewee attempts to deceive the interviewer on selected responses.
10 papers · 0 benchmarks
DiDi (Distractor Distilled Dataset)
DiDi is a distractor-distilled tracking dataset created to address the limitation of low distractor presence in current visual object tracking benchmarks.
10 papers · 1 benchmark
Dynamic Replica is a synthetic dataset of stereo videos featuring humans and animals in virtual environments.
10 papers · 0 benchmarks
ImageNet-VidVRD dataset contains 1,000 videos selected from ILVSRC2016-VID dataset based on whether the video contains clear visual relations.
10 papers · 2 benchmarks
MUGEN is a large-scale video-audio-text dataset MUGEN, collected using the open-sourced platform game CoinRun.
10 papers · 0 benchmarks
Perception Test is a benchmark designed to evaluate the perception and reasoning skills of multimodal models.
10 papers · 3 benchmarks
RVSD (Realistic Video DeSnowing Dataset)
Realistic Video DeSnowing Dataset (RVSD) contains a total of 110 pairs of videos.
10 papers · 0 benchmarks
TAPOS is a new dataset developed on sport videos with manual annotations of sub-actions, and conduct a study on temporal action parsing on top.
10 papers · 1 benchmark
USF (Human ID Gait Challenge Dataset)
The USF Human ID Gait Challenge Dataset is a dataset of videos for gait recognition.
10 papers · 0 benchmarks
We propose a new, scalable video-mining pipeline which transfers captioning supervision from image datasets to video and audio.
10 papers · 0 benchmarks
Composed of 10,000 videos annotated with memorability scores.
10 papers · 0 benchmarks
The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most advanced state-of-the-art methods.
9 papers · 0 benchmarks
BLVD is a large scale 5D semantics dataset collected by the Visual Cognitive Computing and Intelligent Vehicles Lab.
9 papers · 0 benchmarks
The Blackbird unmanned aerial vehicle (UAV) dataset is a large-scale, aggressive indoor flight dataset collected using a custom-built quadrotor platform for use in evaluation of agile perception.
9 papers · 0 benchmarks
LRW-1000 has been renamed as CAS-VSR-W1k.
9 papers · 1 benchmark
We provide manual annotations of 14 semantic keypoints for 100,000 car instances (sedan, suv, bus, and truck) from 53,000 images captured from 18 moving cameras at Multiple intersections in Pittsburgh, PA.
9 papers · 2 benchmarks
The Collective Activity Dataset contains 5 different collective activities: crossing, walking, waiting, talking, and queueing and 44 short video sequences some of which were recorded by consumer hand-held digital camera with varying view…
9 papers · 1 benchmark
EgoHOS (Fine-Grained Egocentric Hand-Object Segmentation Dataset)
EgoHOS is a labeled dataset consisting of 11243 egocentric images with per-pixel segmentation labels of hands and objects being interacted with during a diverse array of daily activities.
9 papers · 0 benchmarks
EgoProceL is a large-scale dataset for procedure learning.
9 papers · 0 benchmarks
FAIR-Play is a video-audio dataset consisting of 1,871 video clips and their corresponding binaural audio clips recording in a music room.
9 papers · 0 benchmarks
HQ-WMCA (High-Quality Wide Multi-Channel Attack database)
The High-Quality Wide Multi-Channel Attack database (HQ-WMCA) database consists of 2904 short multi-modal video recordings of both bona-fide and presentation attacks.
9 papers · 0 benchmarks
Home Action Genome is a large-scale multi-view video database of indoor daily activities.
9 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.