Home › Datasets › modality › Videos

Videos datasets

archive 2025-07-28

1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 1 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Videos datasets 1–48 of 1,014

The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and captioning dataset.
11,922 papers · 77 benchmarks
UCF101 (UCF101 Human Actions dataset)
UCF101 dataset is an extension of UCF50 and consists of 13,320 video clips, which are classified into 101 categories.
1,863 papers · 23 benchmarks
Kinetics (Kinetics Human Action Video Dataset)
The Kinetics dataset is a large-scale, high-quality dataset for human action recognition in videos.
1,341 papers · 18 benchmarks
The HMDB51 dataset is a large collection of realistic videos from various sources, including movies and web videos.
839 papers · 10 benchmarks
The ActivityNet dataset contains 200 different types of activities and a total of 849 hours of videos collected from YouTube.
807 papers · 17 benchmarks
The Human3.6M dataset is one of the largest motion capture datasets, which consists of 3.6 million human poses and corresponding images captured by a high-speed motion capture system.
783 papers · 13 benchmarks
IEMOCAP (The Interactive Emotional Dyadic Motion Capture (IEMOCAP) Database)
Multimodal Emotion Recognition IEMOCAP The IEMOCAP dataset consists of 151 videos of recorded dialogues, with 2 speakers per session for a total of 302 videos across the dataset.
749 papers · 3 benchmarks
Audioset is an audio event dataset, which consists of over 2M human-annotated 10-second video clips.
744 papers · 5 benchmarks
DAVIS (Densely Annotated VIdeo Segmentation)
The Densely Annotation Video Segmentation dataset (DAVIS) is a high quality and high resolution densely annotated video segmentation dataset under two resolutions, 480p and 1080p.
734 papers · 10 benchmarks
MSR-VTT (Microsoft Research Video to Text) is a large-scale dataset for the open domain video captioning, which consists of 10,000 video clips from 20 categories, and each video clip is annotated with 20 English sentences by Amazon…
640 papers · 8 benchmarks
VoxCeleb2 is a large scale speaker recognition dataset obtained automatically from open-source media.
564 papers · 5 benchmarks
Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving.
469 papers · 16 benchmarks
The Charades dataset is composed of 9,848 videos of daily indoors activities with an average length of 30 seconds, involving interactions with 46 objects classes in 15 types of indoor scenes and containing a vocabulary of 30 verbs leading…
428 papers · 6 benchmarks
Object Tracking Benchmark (OTB) is a visual tracking benchmark that is widely used to evaluate the performance of a visual tracking algorithm.
416 papers · 1 benchmark
The 3D Poses in the Wild dataset is the first dataset in the wild with accurate 3D poses for evaluation.
395 papers · 4 benchmarks
The GoPro dataset for deblurring consists of 3,214 blurred images with the size of 1,280×720 that are divided into 2,103 training images and 1,111 test images.
390 papers · 4 benchmarks
FaceForensics++ is a forensics dataset consisting of 1000 original video sequences that have been manipulated with four automated face manipulation methods: Deepfakes, Face2Face, FaceSwap and NeuralTextures.
368 papers · 2 benchmarks
MSVD (Microsoft Research Video Description Corpus)
The Microsoft Research Video Description Corpus (MSVD) dataset consists of about 120K sentences collected during the summer of 2010.
327 papers · 3 benchmarks
The THUMOS14 (THUMOS 2014) dataset is a large-scale video dataset that includes 1,010 videos for validation and 1,574 videos for testing from 20 classes.
318 papers · 18 benchmarks
MOT17 (Multiple Object Tracking 17)
The Multiple Object Tracking 17 (MOT17) dataset is a dataset for multiple object tracking.
291 papers · 2 benchmarks
The 20BN-SOMETHING-SOMETHING V2 dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects.
290 papers · 7 benchmarks
MELD (Multimodal EmotionLines Dataset)
Multimodal EmotionLines Dataset (MELD) has been created by enhancing and extending EmotionLines dataset.
289 papers · 3 benchmarks
MPI-INF-3DHP is a 3D human body pose estimation dataset consisting of both constrained indoor and complex outdoor scenes.
289 papers · 5 benchmarks
HowTo100M is a large-scale dataset of narrated videos with an emphasis on instructional videos where content creators teach complex tasks with an explicit intention of explaining the visual content on screen.
286 papers · 1 benchmark
KTH (KTH Action dataset)
The efforts to create a non-trivial and publicly available dataset for action recognition was initiated at the KTH Royal Institute of Technology in 2004.
279 papers · 2 benchmarks
LaSOT (Large-scale Single Object Tracking)
LaSOT is a high-quality benchmark for Large-scale Single Object Tracking.
275 papers · 3 benchmarks
MSMT17 (Multi Scene Multi Time dataset for person re-id)
MSMT17 is a multi-scene multi-time person re-identification dataset.
275 papers · 6 benchmarks
WebVid contains 10 million video clips with captions, sourced from the web.
257 papers · 1 benchmark
The ActivityNet Captions dataset is built on ActivityNet v1.3 which includes 20k YouTube untrimmed videos with 100k caption annotations.
255 papers · 6 benchmarks
100DOH (100 Days Of Hands Dataset)
The 100 Days Of Hands Dataset (100DOH) is a large-scale video dataset containing hands and hand-object interactions.
249 papers · 0 benchmarks
JHMDB (Joint-annotated Human Motion Data Base)
JHMDB is an action recognition dataset that consists of 960 video sequences belonging to 21 actions.
249 papers · 9 benchmarks
IJB-C (IARPA Janus Benchmark-C)
The IJB-C dataset is a video-based face recognition dataset.
246 papers · 3 benchmarks
GOT-10k (Generic Object Tracking Benchmark)
The GOT-10k dataset contains more than 10,000 video segments of real-world moving objects and over 1.5 million manually labelled bounding boxes.
239 papers · 2 benchmarks
CK+ (Extended Cohn-Kanade dataset)
The Extended Cohn-Kanade (CK+) dataset contains 593 video sequences from a total of 123 different subjects, ranging from 18 to 50 years of age with a variety of genders and heritage.
238 papers · 2 benchmarks
Charades-STA is a new dataset built on top of Charades by adding sentence temporal annotations.
236 papers · 4 benchmarks
DAVIS16 is a dataset for video object segmentation which consists of 50 videos in total (30 videos for training and 20 for testing).
231 papers · 4 benchmarks
CamVid (Cambridge-driving Labeled Video Database)
CamVid (Cambridge-driving Labeled Video Database) is a road/driving scene understanding database which was originally captured as five video sequences with a 960×720 resolution camera mounted on the dashboard of a car.
227 papers · 4 benchmarks
The Vimeo-90K is a large-scale high-quality video dataset for lower-level video processing.
220 papers · 3 benchmarks
DiDeMo (Distinct Describable Moments)
The Distinct Describable Moments (DiDeMo) dataset is one of the largest and most diverse datasets for the temporal localization of events in videos given natural language descriptions.
216 papers · 3 benchmarks
TrackingNet is a large-scale tracking dataset consisting of videos in the wild.
210 papers · 2 benchmarks
The ShanghaiTech Campus dataset has 13 scenes with complex light conditions and camera angles.
207 papers · 4 benchmarks
YouTube-VOS 2018 (Youtube Video Object Segmentation)
Youtube-VOS is a Video Object Segmentation dataset that contains 4,453 videos - 3,471 for training, 474 for validation, and 508 for testing.
203 papers · 10 benchmarks
YouCook2 is the largest task-oriented, instructional video dataset in the vision community.
198 papers · 7 benchmarks
The Moving MNIST dataset contains 10,000 video sequences, each consisting of 20 frames.
194 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.