Home › Datasets › modality › Videos
Videos datasets
archive 2025-07-28
1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 2 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Videos datasets 49–96 of 1,014
The MOTChallenge datasets are designed for the task of multiple object tracking.
192 papers · 0 benchmarks
CMU Multimodal Opinion Sentiment and Emotion Intensity (CMU-MOSEI) is the largest dataset of sentence-level sentiment analysis and emotion recognition in online videos.
190 papers · 3 benchmarks
LRW (Lip Reading in the Wild)
The Lip Reading in the Wild (LRW) dataset a large-scale audio-visual database that contains 500 different words from over 1,000 speakers.
188 papers · 8 benchmarks
OTB-2015, also referred as Visual Tracker Benchmark, is a visual tracking dataset.
182 papers · 1 benchmark
MARS (Motion Analysis and Re-identification Set)
MARS (Motion Analysis and Re-identification Set) is a large scale video based person reidentification dataset, an extension of the Market-1501 dataset.
181 papers · 2 benchmarks
The Breakfast Actions Dataset comprises of 10 actions related to breakfast preparation, performed by 52 different individuals in 18 different kitchens.
179 papers · 6 benchmarks
NExT-QA is a VideoQA benchmark targeting the explanation of video contents.
174 papers · 3 benchmarks
ALFRED (Action Learning From Realistic Environments and Directives)
ALFRED (Action Learning From Realistic Environments and Directives), is a new benchmark for learning a mapping from natural language instructions and egocentric vision to sequences of actions for household tasks.
173 papers · 0 benchmarks
The UCY dataset consist of real pedestrian trajectories with rich multi-human interaction scenarios captured at 2.5 Hz (Δt=0.4s).
170 papers · 1 benchmark
The Sports-1M dataset consists of over a million videos from YouTube.
164 papers · 2 benchmarks
The YCB-Video dataset is a large-scale video dataset for 6D object pose estimation.
164 papers · 5 benchmarks
IJB-B (IARPA Janus Benchmark-B)
The IJB-B dataset is a template-based face dataset that contains 1845 subjects with 11,754 images, 55,025 frames and 7,011 videos where a template consists of a varying number of still images and video frames from different sources.
163 papers · 5 benchmarks
YouTubeVIS is a new dataset tailored for tasks like simultaneous detection, segmentation and tracking of object instances in videos and is collected based on the current largest video object segmentation dataset YouTubeVOS.
163 papers · 2 benchmarks
This paper introduces the pipeline to scale the largest dataset in egocentric vision EPIC-KITCHENS.
162 papers · 6 benchmarks
Video-MME stands for Video Multi-Modal Evaluation.
152 papers · 2 benchmarks
MOT16 (Multiple Object Tracking 2016)
The MOT16 dataset is a dataset for multiple object tracking.
149 papers · 2 benchmarks
DISFA (Denver Intensity of Spontaneous Facial Action)
The Denver Intensity of Spontaneous Facial Action (DISFA) dataset consists of 27 videos of 4844 frames each, with 130,788 images in total.
148 papers · 3 benchmarks
The Kinetics-600 is a large-scale action recognition dataset which consists of around 480K videos from 600 action categories.
148 papers · 3 benchmarks
The YouTube-8M dataset is a large scale video dataset, which includes more than 7 million videos with 4716 classes labeled by the annotation system.
147 papers · 2 benchmarks
The ActivityNet-QA dataset contains 58,000 human-annotated QA pairs on 5,800 videos derived from the popular ActivityNet dataset.
146 papers · 2 benchmarks
The SumMe dataset is a video summarization dataset consisting of 25 videos, each annotated with at least 15 human summaries (390 in total).
146 papers · 3 benchmarks
The TVQA dataset is a large-scale video dataset for video question answering.
146 papers · 3 benchmarks
MPII Human Pose Dataset is a dataset for human pose estimation.
143 papers · 1 benchmark
The UCF-Crime dataset is a large-scale dataset of 128 hours of videos.
142 papers · 3 benchmarks
NTU RGB+D 120 is a large-scale dataset for RGB+D human action recognition, which is collected from 106 distinct subjects and contains more than 114 thousand video samples and 8 million frames.
137 papers · 8 benchmarks
SEED-Bench consists of 19K multiple choice questions with accurate human annotations (~6 larger than existing benchmarks), which spans 12 evaluation dimensions including the comprehension of both the image and video modality.
137 papers · 0 benchmarks
Cholec80 is an endoscopic video dataset containing 80 videos of cholecystectomy surgeries performed by 13 surgeons.
134 papers · 2 benchmarks
The Replay-Attack Database for face spoofing consists of 1300 video clips of photo and video attack attempts to 50 clients, under different lighting conditions.
133 papers · 1 benchmark
Virtual KITTI is a photo-realistic synthetic video dataset designed to learn and evaluate computer vision models for several video understanding tasks: object detection and multi-object tracking, scene-level and instance-level semantic…
133 papers · 0 benchmarks
VOT2018 is a dataset for visual object tracking.
129 papers · 1 benchmark
DFDC (Deepfake Detection Challenge)
The DFDC (Deepfake Detection Challenge) is a dataset for deepface detection consisting of more than 100,000 videos.
126 papers · 1 benchmark
FBMS (Freiburg-Berkeley Motion Segmentation)
The Freiburg-Berkeley Motion Segmentation Dataset (FBMS-59) is an extension of the BMS dataset with 33 additional video sequences.
126 papers · 1 benchmark
LSMDC (Large Scale Movie Description Challenge)
This dataset contains 118,081 short video clips extracted from 202 movies.
126 papers · 3 benchmarks
SUN3D contains a large-scale RGB-D video database, with 8 annotated sequences.
126 papers · 0 benchmarks
GTEA (Georgia Tech Egocentric Activity)
The Georgia Tech Egocentric Activities (GTEA) dataset contains seven types of daily activities such as making sandwich, tea, or coffee.
120 papers · 2 benchmarks
VATEX is multilingual, large, linguistically complex, and diverse dataset in terms of both video and natural language descriptions.
118 papers · 3 benchmarks
The KVASIR Dataset was released as part of the medical multimedia challenge presented by MediaEval.
117 papers · 1 benchmark
The 20BN-SOMETHING-SOMETHING dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects.
117 papers · 3 benchmarks
LRS2 (Lip Reading Sentences 2)
The Oxford-BBC Lip Reading Sentences 2 (LRS2) dataset is one of the largest publicly available datasets for lip reading sentences in-the-wild.
115 papers · 10 benchmarks
Over a period of three years (2009 - 2011) the daily news and weather forecast airings of the German public tv-station PHOENIX featuring sign language interpretation have been recorded and the weather forecasts of a subset of 386 editions…
114 papers · 2 benchmarks
AVA (Atomic Visual Actions)
AVA is a project that provides audiovisual annotations of video for improving our understanding of human activity.
113 papers · 7 benchmarks
VOT2016 is a video dataset for visual object tracking.
113 papers · 1 benchmark
EgoSchema is very long-form video question-answering dataset, and benchmark to evaluate long video understanding capabilities of modern vision and language systems.
112 papers · 3 benchmarks
OTB2013 is the previous version of the current OTB2015 Visual Tracker Benchmark.
110 papers · 2 benchmarks
The Penn Action Dataset contains 2326 video sequences of 15 different actions and human joint annotations for each sequence.
110 papers · 4 benchmarks
SegTrack v2 is a video segmentation dataset with full pixel-level annotations on multiple objects at each frame within each video.
107 papers · 5 benchmarks
The COIN dataset (a large-scale dataset for COmprehensive INstructional video analysis) consists of 11,827 videos related to 180 different tasks in 12 domains (e.g., vehicles, gadgets, etc.) related to our daily life.
105 papers · 2 benchmarks
JIGSAWS (JHU-ISI Gesture and Skill Assessment Working Set)
The JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS) is a surgical activity dataset for human motion modeling.
105 papers · 3 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.