Home › Datasets › modality › Videos

Videos datasets

archive 2025-07-28

1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 4 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Videos datasets 145–192 of 1,014

Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles.
57 papers · 4 benchmarks
The EYEDIAP dataset is a dataset for gaze estimation from remote RGB, and RGB-D (standard vision and depth), cameras.
57 papers · 3 benchmarks
The TotalCapture dataset consists of 5 subjects performing several activities such as walking, acting, a range of motion sequence (ROM) and freestyle motions, which are recorded using 8 calibrated, static HD RGB cameras and 13 IMUs…
57 papers · 2 benchmarks
VOT2017 (Visual Object Tracking Challenge)
VOT2017 is a Visual Object Tracking dataset for different tasks that contains 60 short sequences annotated with 6 different attributes.
56 papers · 2 benchmarks
CrossTask dataset contains instructional videos, collected for 83 different tasks.
54 papers · 1 benchmark
MovieNet is a holistic dataset for movie understanding.
54 papers · 1 benchmark
PandaSet is a dataset produced by a complete, high-precision autonomous vehicle sensor kit with a no-cost commercial license.
54 papers · 0 benchmarks
RxR (Room-across-Room)
Room-Across-Room (RxR) is a multilingual dataset for Vision-and-Language Navigation (VLN) for Matterport3D environments.
54 papers · 1 benchmark
The SEMAINE videos dataset contains spontaneous data capturing the audiovisual interaction between a human and an operator undertaking the role of an avatar with four personalities: Poppy (happy), Obadiah (gloomy), Spike (angry) and…
54 papers · 1 benchmark
Large language models (LLMs), after being aligned with vision models and integrated into vision-language models (VLMs), can bring impressive improvement in image reasoning tasks.
53 papers · 1 benchmark
RICH (Real scenes, Interaction, Contact and Humans)
Inferring human-scene contact (HSC) is the first step toward understanding how humans interact with their surroundings.
53 papers · 1 benchmark
Consists of 100 challenging video sequences captured from real-world traffic scenes (over 140,000 frames with rich annotations, including occlusion, weather, vehicle category, truncation, and vehicle bounding boxes) for object detection,…
53 papers · 2 benchmarks
There exist previous works [6, 10] that constructed referring segmentation datasets for videos.
52 papers · 3 benchmarks
Sprites (2D Video Game Character Sprites)
The Sprites dataset contains 60 pixel color images of animated characters (sprites).
52 papers · 3 benchmarks
YouTube-VIS 2021 (Video Instance Segmentation on YouTube-VIS 2021 validation)
3,859 high-resolution YouTube videos, 2,985 training videos, 421 validation videos and 453 test videos.
52 papers · 1 benchmark
Rendered synthetically using a library of standard 3D objects, and tests the ability to recognize compositions of object movements that require long-term reasoning.
51 papers · 3 benchmarks
The large-scale MUSIC-AVQA dataset of musical performance contains 45,867 question-answer pairs, distributed in 9,288 videos for over 150 hours.
51 papers · 1 benchmark
TGIF (Tumblr GIF)
The Tumblr GIF (TGIF) dataset contains 100K animated GIFs and 120K sentences describing visual content of the animated GIFs.
51 papers · 1 benchmark
MOSE (Complex Video Object Segmentation)
CoMplex video Object SEgmentation (MOSE) is a dataset to study the tracking and segmenting objects in complex environments.
49 papers · 2 benchmarks
TAO (Tracking Any Object Dataset)
TAO is a federated dataset for Tracking Any Object, containing 2,907 high resolution videos, captured in diverse environments, which are half a minute long on average.
49 papers · 1 benchmark
WildDeepfake is a dataset for real-world deepfakes detection which consists of 7,314 face sequences extracted from 707 deepfake videos that are collected completely from the internet.
49 papers · 0 benchmarks
CityFlow is a city-scale traffic camera dataset consisting of more than 3 hours of synchronized HD videos from 40 cameras across 10 intersections, with the longest distance between two simultaneous cameras being 2.5 km.
47 papers · 1 benchmark
FLIC (Frames Labelled in Cinema)
The FLIC dataset contains 5003 images from popular Hollywood movies.
46 papers · 2 benchmarks
InternVid is a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodAL understanding and generation.
46 papers · 0 benchmarks
OpenLane is the first real-world and the largest scaled 3D lane dataset to date.
46 papers · 2 benchmarks
TAP-Vid is a benchmark which contains both real-world videos with accurate human annotations of point tracks, and synthetic videos with perfect ground-truth point tracks.
46 papers · 1 benchmark
BEDLAM is a large-scale synthetic video dataset designed to train and test algorithms on the task of 3D human pose and shape estimation (HPS).
45 papers · 1 benchmark
DailyActivity3D dataset is a daily activity dataset captured by a Kinect device.
45 papers · 1 benchmark
This data set was prepared from 88 open-source YouTube cooking videos.
45 papers · 0 benchmarks
The dataset contains over 15K images of 20 people (6 females and 14 males - 4 people were recorded twice).
44 papers · 1 benchmark
A2D (Actor-Action Dataset)
A2D (Actor-Action Dataset) is a dataset for simultaneously inferring actors and actions in videos.
42 papers · 1 benchmark
AVSpeech is a large-scale audio-visual dataset comprising speech clips with no interfering background signals.
42 papers · 0 benchmarks
DeformingThings4D is a synthetic dataset containing 1,972 animation sequences spanning 31 categories of humanoids and animals.
42 papers · 0 benchmarks
The EPIC-KITCHENS-55 dataset comprises a set of 432 egocentric videos recorded by 32 participants in their kitchens at 60fps with a head mounted camera.
42 papers · 3 benchmarks
ACID (Aerial Coastline Imagery Dataset)
ACID consists of thousands of aerial drone videos of different coastline and nature scenes on YouTube.
41 papers · 1 benchmark
NT-VOT211 consists of 211 diverse videos, offering 211,000 well-annotated frames with 8 attributes including camera motion, deformation, fast motion, motion blur, tiny target, distractors, occlusion and out-of-view.
41 papers · 1 benchmark
QVHighlights (Query-based Video Highlights)
The Query-based Video Highlights (QVHighlights) dataset is a dataset for detecting customized moments and highlights from videos given natural language (NL).
41 papers · 4 benchmarks
The RGBT234 dataset is a comprehensive video dataset specifically designed for RGB-T (Red-Green-Blue and Thermal) tracking purposes.
38 papers · 1 benchmark
URMP (University of Rochester Multi-Modal Musical Performance)
URMP (University of Rochester Multi-Modal Musical Performance) is a dataset for facilitating audio-visual analysis of musical performances.
38 papers · 2 benchmarks
The EgoGesture dataset contains 2,081 RGB-D videos, 24,161 gesture samples and 2,953,224 frames from 50 distinct subjects.
37 papers · 2 benchmarks
CelebV-HQ is a large-scale video facial attributes dataset with annotations.
36 papers · 3 benchmarks
JTA (Joint Track Auto)
JTA is a dataset for people tracking in urban scenarios by exploiting a photorealistic videogame.
36 papers · 1 benchmark
MAD (Movie Audio Descriptions) is an automatically curated large-scale dataset for the task of natural language grounding in videos or natural language moment retrieval.
36 papers · 2 benchmarks
Activity recognition research has shifted focus from distinguishing full-body motion patterns to recognizing complex interactions of multiple entities.
35 papers · 2 benchmarks
The BirdSong dataset consists of audio recordings of bird songs at the H.
35 papers · 0 benchmarks
V2V4Real is a large-scale real-world multi-modal dataset for V2V perception.
35 papers · 0 benchmarks
The EgoHands dataset contains 48 Google Glass videos of complex, first-person interactions between two people.
34 papers · 0 benchmarks
MOT20 is a dataset for multiple object tracking.
34 papers · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.