Home › Datasets › modality › Videos

Videos datasets

archive 2025-07-28

1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 6 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Videos datasets 241–288 of 1,014

Toyota Smarthome dataset (Toyota Smarthome Trimmed)
Toyota Smarthome Trimmed has been designed for the activity classification task of 31 activities.
23 papers · 0 benchmarks
A3D (AnAn Accident Detection)
A new dataset of diverse traffic accidents.
22 papers · 1 benchmark
The Easy Communications (EasyCom) dataset is a world-first dataset designed to help mitigate the cocktail party effect from an augmented-reality (AR) -motivated multi-sensor egocentric world view.
22 papers · 4 benchmarks
TVBench is a new benchmark specifically created to evaluate temporal understanding in video QA.
22 papers · 1 benchmark
iVQA (Instructional Video Question Answering)
An open-ended VideoQA benchmark that aims to: i) provide a well-defined evaluation by including five correct answer annotations per question and ii) avoid questions which can be answered without the video.
22 papers · 2 benchmarks
AIST++ is a 3D dance dataset which contains 3D motion reconstructed from real dancers paired with music.
21 papers · 2 benchmarks
JAAD (Joint Attention in Autonomous Driving)
JAAD is a dataset for studying joint attention in the context of autonomous driving.
21 papers · 1 benchmark
The UT-Interaction dataset contains videos of continuous executions of 6 classes of human-human interactions: shake-hands, point, hug, push, kick and punch.
21 papers · 1 benchmark
Consists of 1106 action samples from seven actions with quality scores as measured by expert human judges.
20 papers · 1 benchmark
AVSBench (Audio −Visual Segmentation)
AVSBench is a pixel-level audio-visual segmentation benchmark that provides ground truth labels for sounding objects.
20 papers · 0 benchmarks
MSU NR VQA Database (MSU No-Reference Video Quality Assessment Database)
The dataset was created for video quality assessment problem.
20 papers · 2 benchmarks
Spatio-temporal action detection is an important and challenging problem in video understanding.
20 papers · 2 benchmarks
The iLIDS-VID dataset is a person re-identification dataset which involves 300 different pedestrians observed across two disjoint camera views in public open space.
20 papers · 2 benchmarks
Argoverse-HD is a dataset built for streaming object detection, which encompasses real-time object detection, video object detection, tracking, and short-term forecasting.
19 papers · 4 benchmarks
CMD (Condensed Movies Dataset)
Consists of the key scenes from over 3K movies: each key scene is accompanied by a high level semantic description of the scene, character face-tracks, and metadata about the movie.
19 papers · 0 benchmarks
FBMS-59 (Freiburg-Berkeley Motion Segmentation)
The Freiburg-Berkeley Motion Segmentation Dataset (FBMS-59) is a dataset for motion segmentation, which extends the BMS-26 dataset with 33 additional video sequences.
19 papers · 3 benchmarks
HiEve (Human-in-Events)
A new large-scale dataset for understanding human motions, poses, and actions in a variety of realistic events, especially crowd & complex events.
19 papers · 1 benchmark
LVOS is a dataset for long-term video object segmentation (VOS).
19 papers · 0 benchmarks
The MECCANO dataset is the first dataset of egocentric videos to study human-object interactions in industrial-like settings.
19 papers · 3 benchmarks
SUTD-TrafficQA (Singapore University of Technology and Design - Traffic Question Answering) is a dataset which takes the form of video QA based on 10,080 in-the-wild videos and annotated 62,535 QA pairs, for benchmarking the cognitive…
19 papers · 1 benchmark
A dataset with fully annotated attention targets in video for attention target estimation.
19 papers · 1 benchmark
WanJuan is a large-scale training corpus that includes multiple modalities.
19 papers · 0 benchmarks
ActivityNet-Entities, augments the challenging ActivityNet Captions dataset with 158k bounding box annotations, each grounding a noun phrase.
18 papers · 0 benchmarks
BOBSL (BBC-Oxford British Sign Language)
BOBSL is a large-scale dataset of British Sign Language (BSL).
18 papers · 1 benchmark
CASIA-B is a large multiview gait database, which is created in January 2005.
18 papers · 1 benchmark
CCD (Car Crash Dataset)
Car Crash Dataset (CCD) is collected for traffic accident analysis.
18 papers · 1 benchmark
FDST (Fudan-ShanghaiTech)
The Fudan-ShanghaiTech dataset (FDST) is a dataset for video crowd counting.
18 papers · 0 benchmarks
KoNViD-1k (KoNViD-1k VQA Database)
Subjective video quality assessment (VQA) strongly depends on semantics, context, and the types of visual distortions.
18 papers · 1 benchmark
MSU FR VQA Database (MSU Full-Reference Video Quality Assessment Database)
The dataset was created for video quality assessment problem.
18 papers · 2 benchmarks
We create a benchmark dataset named ReVOS.
18 papers · 1 benchmark
The Replay-Mobile Database for face spoofing consists of 1190 video clips of photo and video attack attempts to 40 clients, under different lighting conditions.
18 papers · 0 benchmarks
SCAND (Socially CompliAnt Navigation Dataset)
Have you wondered how autonomous mobile robots should share space with humans in public spaces?
18 papers · 0 benchmarks
SeaDronesSee (SeaDronesSee: A Maritime Benchmark for Detecting Humans in Open Water)
SeaDronesSee is a large-scale data set aimed at helping develop systems for Search and Rescue (SAR) using Unmanned Aerial Vehicles (UAVs) in maritime scenarios.
18 papers · 3 benchmarks
VidSitu is a dataset for the task of semantic role labeling in videos (VidSRL).
18 papers · 0 benchmarks
DAiSEE is a multi-label video classification dataset comprising of 9,068 video snippets captured from 112 users for recognizing the user affective states of boredom, confusion, engagement, and frustration "in the wild".
17 papers · 1 benchmark
EgoTask QA benchmark contains 40K balanced question-answer pairs selected from 368K programmatically generated questions generated over 2K egocentric videos.
17 papers · 1 benchmark
Memorability dataset with 10000 3-second videos.
17 papers · 0 benchmarks
MoCA-Mask (Moving Camouflaged Animals (MoCA)-Mask)
The original Moving Camouflaged Animals (MoCA) Dataset includes 37K frames from 141 YouTube Video sequences with resolution and sampling rate of 720 × 1280 and 24fps in the majority of cases.
17 papers · 1 benchmark
OMG-Emotion (One-Minute Gradual-Emotional Behavior)
The One-Minute Gradual-Emotional Behavior dataset (OMG-Emotion) dataset is composed of Youtube videos which are around a minute in length and are annotated taking into consideration a continuous emotional behavior.
17 papers · 0 benchmarks
STAR Benchmark (Situated Reasoning)
How to capture the present knowledge from surrounding situations and perform reasoning accordingly is crucial and challenging for machine intelligence.
17 papers · 2 benchmarks
The SUN-SEG dataset is a high-quality per-frame annotated VPS dataset, which includes 158,690 frames from the famous SUN dataset.
17 papers · 1 benchmark
The SUN-SEG dataset is a high-quality per-frame annotated VPS dataset, which includes 158,690 frames from the famous SUN dataset.
17 papers · 1 benchmark
VIPL-HR database is a database for remote heart rate (HR) estimation from face videos under less-constrained situations.
17 papers · 1 benchmark
A temporal counterfactual dataset composing of 1000 short and natural video-caption pairs.
17 papers · 1 benchmark
CMU-MOSI (Multimodal Corpus of Sentiment Intensity)
The Multimodal Corpus of Sentiment Intensity (CMU-MOSI) dataset is a collection of 2199 opinion video clips.
16 papers · 2 benchmarks
Casual Conversations dataset is designed to help researchers evaluate their computer vision and audio models for accuracy across a diverse set of age, genders, apparent skin tones and ambient lighting conditions.
16 papers · 0 benchmarks
CholecT45 is a subset of CholecT50 consisting of 45 videos from the Cholec80 dataset.
16 papers · 1 benchmark
Expi (Extreme Pose Interaction)
Extreme Pose Interaction (ExPI) Dataset is a new person interaction dataset of Lindy Hop dancing actions.
16 papers · 3 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.