Home › Datasets › modality › Videos

Videos datasets

archive 2025-07-28

1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 15 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Videos datasets 673–720 of 1,014

Large-scale Anomaly Detection (LAD) is a database to benchmark anomaly detection in video sequences, which is featured in two aspects.
2 papers · 0 benchmarks
This data set contains 775 video sequences, captured in the wildlife park Lindenthal (Cologne, Germany) as part of the AMMOD project, using an Intel RealSense D435 stereo camera.
2 papers · 0 benchmarks
LoTE-Animal (LoTE-Animal: A Long Time-span Dataset for Endangered Animal Behavior Understanding)
Understanding and analyzing animal behavior is increasingly essential to protect endangered animal species.
2 papers · 1 benchmark
Despite impressive advancements in video understanding, most efforts remain limited to coarse-grained or visual-only video tasks.
2 papers · 0 benchmarks
Lyra Dataset (A Dataset for Greek Traditional and Folk Music)
Lyra is a dataset of 1570 traditional and folk Greek music pieces that includes audio and video (timestamps and links to YouTube videos), along with annotations that describe aspects of particular interest for this dataset, including…
2 papers · 0 benchmarks
The Model for Attended Awareness in Driving (MAAD) is a dataset of third-person estimates of a driver’s attended awareness.
2 papers · 0 benchmarks
MAVS (Multilingual Audio-Visual Smartphone dataset)
MAVS is an audio-visual smartphone dataset captured in five different recent smartphones.
2 papers · 0 benchmarks
MCAD (Multi-Camera Action Dataset)
Designed to evaluate the open view classification problem under the surveillance environment.
2 papers · 0 benchmarks
MECD (Multi-Event Causal Discovery)
Provide: 1,105 lifestyle videos that span diverse scenarios.
2 papers · 1 benchmark
We provide a dataset called MMAC Captions for sensor-augmented egocentric-video captioning.
2 papers · 0 benchmarks
MetaVD (Meta Video Dataset)
MetaVD is a Meta Video Dataset for enhancing human action recognition datasets.
2 papers · 0 benchmarks
MoB (Malicious or Benign Cartoon Videos)
A dataset of cartoon video clips.
2 papers · 1 benchmark
MuCo-VQA consist of large-scale (3.7M) multilingual and code-mixed VQA datasets in multiple languages: Hindi (hi), Bengali (bn), Spanish (es), German (de), French (fr) and code-mixed language pairs: en-hi, en-bn, en-fr, en-de and en-es.
2 papers · 0 benchmarks
MultiOOD (Multimodal Out-of-Distribution Detection Benchmark)
MultiOOD is the first benchmark for Multimodal OOD Detection and covers diverse dataset sizes and modalities.
2 papers · 0 benchmarks
NSVA (NBA dataset for Sports Video Analysis)
NVSA is a large-scale NBA dataset for Sports Video Analysis (NSVA) with a focus on sports video captioning.
2 papers · 0 benchmarks
A new video dataset for OR, with 30, 000 objects over 5, 000 stereo video sequences annotated for their descriptions and gaze.
2 papers · 0 benchmarks
PESMOD (PExels Small Moving Object Detection)
The PESMOD (PExels Small Moving Object Detection) dataset consists of high resolution aerial images in which moving objects are labelled manually.
2 papers · 0 benchmarks
The data includes all movement trajectories extracted from the videos of Parkinson's assessments using Convolutional Pose Machines (CPM) as well as the confidence values from CPM.
2 papers · 0 benchmarks
A set of 221 stereo videos captured by the SOCRATES stereo camera trap in a wildlife park in Bonn, Germany between February and July of 2022.
2 papers · 0 benchmarks
R2VQ (Recipe-to-Video Questions)
R2VQ is a dataset designed for testing competence-based comprehension of machines over a multimodal recipe collection, which contains text-video aligned recipes.
2 papers · 0 benchmarks
RGBD1K (A Large-scale Dataset and Benchmark for RGB-D Object Tracking)
RGBD1K is a benchmark for RGB-D Object Tracking which contains 1050 sequences with about 2.5M frames in total.
2 papers · 0 benchmarks
RHM (Rhm: Robot house multi-view human activity recognition dataset)
The Robot House Multi-View dataset (RHM) contains four views: Front, Back, Ceiling, and Robot Views.
2 papers · 1 benchmark
RISEdb (Robust Indoor Localization in Complex Scenarios (RISE) database)
The RISE (Robust Indoor Localization in Complex Scenarios) dataset is meant to train and evaluate visual indoor place recognizers.
2 papers · 0 benchmarks
RLD (Responsive Listener Dataset)
RLD (Responsive Listener Dataset) is a conversation video corpus collected from the public resources featuring 67 speakers, 76 listeners with three different attitudes.
2 papers · 0 benchmarks
The Rhythmic Gymnastics dataset contains videos of four different types of gymnastics routines: ball, clubs, hoop and ribbon.
2 papers · 1 benchmark
SB20 (Sugar Beet 2020 University of Bonn)
Video sequences captured at a field on Campus Kleinaltendorf (CKA), University of Bonn, captured by BonBot-I, an autonomous weeding robot.
2 papers · 0 benchmarks
SEPE 8K dataset is made of 40 different 8K (8192 x 4320) video sequences and 40 variant 8K (8192 x 5464) images.
2 papers · 1 benchmark
Sakuga-42M is a large-scale hand-drawn cartoon video dataset for academic research purposes, it comprises 42 million cartoon keyframes covering various artistic styles, regions, and years, with comprehensive semantic annotations including…
2 papers · 0 benchmarks
A curated and 3-D pose-annotated subset of RGB videos sourced from Kinetics-700, a large-scale action dataset.
2 papers · 1 benchmark
A dataset derived from the recently introduced Mimetics dataset.
2 papers · 2 benchmarks
Stanford-ECM is an egocentric multimodal dataset which comprises about 27 hours of egocentric video augmented with heart rate and acceleration data.
2 papers · 0 benchmarks
StoryBench (StoryBench: A Multifaceted Benchmark for Continuous Story Visualization)
StoryBench is a multi-task benchmark to reliably evaluate the ability of text-to-video models to generate stories from a sequence of captions and their duration.
2 papers · 1 benchmark
SuHiFiMask (Surveillance High-Fidelity Mask)
SuHiFiMask (Surveillance High-Fidelity Mask) extends FAS to real surveillance scenes rather than mimicking low-resolution images and surveillance environments.
2 papers · 0 benchmarks
SynthEVox3D-Tiny (Synthetic Event Camera Voxel 3D Reconstruction Dataset)
Event cameras are sensors that are inspired by biological systems and specialize in capturing changes in brightness.
2 papers · 1 benchmark
TAPVid-3D is a dataset and benchmark for evaluating the task of long-range Tracking Any Point in 3D (TAP-3D).
2 papers · 0 benchmarks
TVL Dataset (Touch-Vision-Language Dataset)
Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model.
2 papers · 0 benchmarks
ThermoHands is the first benchmark dataset specifically designed for egocentric 3D hand pose estimation from thermal images.
2 papers · 0 benchmarks
Twitch-FIFA is video-context, many-speaker dialogue dataset based on live-broadcast soccer game videos and chats from Twitch.tv.
2 papers · 0 benchmarks
The Udacity dataset is mainly composed of video frames taken from urban roads.
2 papers · 1 benchmark
V2VBench is a comprehensive benchmark designed to evaluate video editing methods.
2 papers · 0 benchmarks
VGG-Sound Sync is an audio-visual synchronisation benchmark based on videos collected from YouTube.
2 papers · 0 benchmarks
VTC (Videos, Titles and Comments)
VTC is a large-scale multimodal dataset containing video-caption pairs (~300k) alongside comments that can be used for multimodal representation learning.
2 papers · 0 benchmarks
VideoForensicsHQ is a benchmark dataset for face video forgery detection, providing high quality visual manipulations.
2 papers · 0 benchmarks
We present YTSeg, a topically and structurally diverse benchmark for the text segmentation task based on YouTube transcriptions.
2 papers · 1 benchmark
The YouTube8M-MusicTextClips dataset consists of over 4k high-quality human text descriptions of music found in video clips from the YouTube8M dataset.
2 papers · 0 benchmarks
The dataset consists of high-resolution three-dimensional (3D) turbulent flow simulations.
1 paper · 0 benchmarks
The 42Street dataset is based on a theater play as an example of such an application.
1 paper · 0 benchmarks
A Ball-Collision Dataset (ABCD) serves as a comprehensive benchmark for investigating the interaction dynamics of moving objects within 3D environments.
1 paper · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.