Home › Datasets › modality › Videos

Videos datasets

archive 2025-07-28

1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 3 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Videos datasets 97–144 of 1,014

The BP4D-Spontaneous dataset is a 3D video database of spontaneous facial expressions in a diverse group of young adults.
104 papers · 3 benchmarks
The PoseTrack dataset is a large-scale benchmark for multi-person pose estimation and tracking in videos.
103 papers · 5 benchmarks
CUHK-SYSU (CUHK-SYSU Person Search Dataset)
The CUKL-SYSY dataset is a large scale benchmark for person search, containing 18,184 images and 8,432 identities.
100 papers · 2 benchmarks
UAVDT (Unmanned Aerial Vehicle Benchmark Object Detection and Tracking)
UAVDT is a large scale challenging UAV Detection and Tracking benchmark (i.e., about 80, 000 representative frames from 10 hours raw videos) for 3 important fundamental tasks, i.e., object DETection (DET), Single Object Tracking (SOT) and…
96 papers · 2 benchmarks
Kinetics-700 is a video dataset of 650,000 clips that covers 700 human action classes.
95 papers · 3 benchmarks
DexYCB is a dataset for capturing hand grasping of objects.
94 papers · 2 benchmarks
The Hopkins 155 dataset consists of 156 video sequences of two or three motions.
92 papers · 1 benchmark
The TGIF-QA dataset contains 165K QA pairs for the animated GIFs from the TGIF dataset [Li et al.
92 papers · 3 benchmarks
UCSD Ped2 (UCSD Anomaly Detection Dataset)
The UCSD Anomaly Detection Dataset was acquired with a stationary camera mounted at an elevation, overlooking pedestrian walkways.
92 papers · 4 benchmarks
A large-scale multi-object tracking dataset for human tracking in occlusion, frequent crossover, uniform appearance and diverse body gestures.
91 papers · 1 benchmark
The MSU-MFSD dataset contains 280 video recordings of genuine and attack faces.
89 papers · 1 benchmark
FaceForensics is a video dataset consisting of more than 500,000 frames containing faces from 1004 videos that can be used to study image or video forgeries.
88 papers · 1 benchmark
The MovieQA dataset is a dataset for movie question answering.
86 papers · 1 benchmark
CAMUS (Cardiac Acquisitions for Multi-structure Ultrasound Segmentation)
This project aims to provide all the materials to the community to resolve the problem of echocardiographic image segmentation and volume estimation from 2D ultrasound sequences (both two and four-chamber views).
85 papers · 0 benchmarks
CULane is a large scale challenging dataset for academic research on traffic lane detection.
85 papers · 1 benchmark
The How2 dataset contains 13,500 videos, or 300 hours of speech, and is split into 185,187 training, 2022 development (dev), and 2361 test utterances.
84 papers · 2 benchmarks
Oulu-CASIA (Oulu-CASIA NIR&VIS facial expression database)
The Oulu-CASIA NIR&VIS facial expression database consists of six expressions (surprise, happiness, sadness, anger, fear and disgust) from 80 people between 23 and 58 years old.
80 papers · 4 benchmarks
Volleyball is a video action recognition dataset.
80 papers · 3 benchmarks
PRW (Person Re-identification in the Wild)
PRW is a large-scale dataset for end-to-end pedestrian detection and person recognition in raw video frames.
77 papers · 1 benchmark
FineGym is an action recognition dataset build on top of gymnasium videos.
76 papers · 0 benchmarks
A2D2 (Audi Autonomous Driving Dataset)
Audi Autonomous Driving Dataset (A2D2) consists of simultaneously recorded images and 3D point clouds, together with 3D bounding boxes, semantic segmentation, instance segmentation, and data extracted from the automotive bus.
75 papers · 0 benchmarks
HACS (Human Action Clips and Segments)
HACS is a dataset for human action recognition.
75 papers · 2 benchmarks
OVIS (Occluded Video Instance Segmentation)
OVIS is a new large scale benchmark dataset for video instance segmentation task.
75 papers · 1 benchmark
ApolloScape is a large dataset consisting of over 140,000 video frames (73 street scene videos) from various locations in China under varying weather conditions.
74 papers · 4 benchmarks
DDAD (Dense Depth for Autonomous Driving)
DDAD is a new autonomous driving benchmark from TRI (Toyota Research Institute) for long range (up to 250m) and dense depth estimation in challenging and diverse urban conditions.
73 papers · 1 benchmark
MSP-IMPROV (MSP-IMPROV: An Acted Corpus of Dyadic Interactions to Study Emotion Perception)
We present the MSP-IMPROV corpus, a multimodal emotional database, where the goal is to have control over lexical content and emotion while also promoting naturalness in the recordings.
70 papers · 1 benchmark
OmniObject3D is a large vocabulary 3D object dataset with massive high-quality real-scanned 3D objects.
70 papers · 0 benchmarks
YouTube-UGC (YouTube UGC dataset)
This YouTube dataset is a sampling from thousands of User Generated Content (UGC) as uploaded to YouTube distributed under the Creative Commons license.
70 papers · 1 benchmark
Multimodal Opinionlevel Sentiment Intensity (MOSI) contains: (1) multimodal observations including transcribed speech and visual gestures as well as automatic audio and visual features, (2) opinion-level subjectivity segmentation, (3)…
69 papers · 1 benchmark
MOT15 (Multiple Object Tracking 15)
MOT2015 is a dataset for multiple object tracking.
67 papers · 5 benchmarks
WLASL (Word-Level American Sign Language)
WLASL is a large video dataset for Word-Level American Sign Language (ASL) recognition, which features 2,000 common different words in ASL.
66 papers · 3 benchmarks
The CAD-60 and CAD-120 data sets comprise of RGB-D video sequences of humans performing activities which are recording using the Microsoft Kinect sensor.
65 papers · 1 benchmark
FakeAVCeleb is a novel Audio-Video Deepfake dataset that not only contains deepfake videos but respective synthesized cloned audios as well.
65 papers · 2 benchmarks
MMI (MMI Facial Expression Database)
The MMI Facial Expression Database consists of over 2900 videos and high-resolution still images of 75 subjects.
65 papers · 1 benchmark
Mall (Mall Dataset)
The Mall is a dataset for crowd counting and profiling research.
64 papers · 1 benchmark
CSL-Daily (Chinese Sign Language Corpus) is a large-scale continuous SLT dataset.
63 papers · 2 benchmarks
LRS3-TED is a multi-modal dataset for visual and audio-visual speech recognition.
63 papers · 7 benchmarks
COCO-QA is a dataset for visual question answering.
62 papers · 0 benchmarks
The UTD-MHAD dataset consists of 27 different actions performed by 8 subjects.
62 papers · 2 benchmarks
SFEW (Static Facial Expression in the Wild)
The Static Facial Expressions in the Wild (SFEW) dataset is a dataset for facial expression recognition.
61 papers · 1 benchmark
TVQA+ contains 310.8K bounding boxes, linking depicted objects to visual concepts in questions and answers.
60 papers · 0 benchmarks
CASIA-MFSD is a dataset for face anti-spoofing.
58 papers · 1 benchmark
The MultiTHUMOS dataset contains dense, multilabel, frame-level action annotations for 30 hours across 400 videos in the THUMOS'14 action detection dataset.
58 papers · 3 benchmarks
SQA3D (Situated Question Answering in 3D Scenes)
SQA3D is a dataset for embodied scene understanding, where an agent needs to comprehend the scene it situates from an first person's perspective and answer questions.
58 papers · 3 benchmarks
A novel large-scale corpus of manual annotations for the SoccerNet video dataset, along with open challenges to encourage more research in soccer understanding and broadcast production.
58 papers · 6 benchmarks
XD-Violence is a large-scale audio-visual dataset for violence detection in videos.
58 papers · 2 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.