Home › Datasets › modality › Videos

Videos datasets

archive 2025-07-28

1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 17 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Videos datasets 769–816 of 1,014

DAVIS-Edit is a curated testing benchmark for video editing.
1 paper · 0 benchmarks
DIO (Discovering Interacted Objects)
Discovering Interacted Objects (DIO) is a benchmark containing 51 interactions and 1,000+ objects designed for Spatio-temporal Human-Object Interaction (ST-HOI) detection.
1 paper · 0 benchmarks
DailyMoth-70h is a fully self-contained ASL-to-English sign language dataset containing over 70h of video (48K clips) with aligned English captions of a single native ASL signer (white, male, and early middle-aged) from the ASL news…
1 paper · 0 benchmarks
The Daimler Monocular Pedestrian Detection dataset is a dataset for pedestrian detection in urban environments.
1 paper · 0 benchmarks
The laparoscopic surgery dataset is associated with our International Journal of Computer Assisted Radiology and Surgery (IJCARS) publication titled “DeSmoke-LAP: Improved Unpaired Image-to-Image Translation for Desmoking in Laparoscopic…
1 paper · 0 benchmarks
DeVAn (Dense Video Annotation for Video-Language Models)
DeVAn is a multi-modal dataset containing 8.5K video clips carefully selected from previously published YouTube-based video datasets (YouTube-8M and YT-Temporal-1B) that integrate visual and auditory information.
1 paper · 0 benchmarks
Demo video (Demonstration of Stickbug Robot)
Demonstration video of the Stickbug Robot
1 paper · 0 benchmarks
Dense Forest Trail is an UAV dataset collected from a variety of simulated environment in Unreal Engine.
1 paper · 0 benchmarks
Driver Micro Hand Gestures (DriverMHG) is a dataset for dynamic recognition of driver micro hand gestures, which consists of RGB, depth and infrared modalities.
1 paper · 0 benchmarks
This dataset consists of a number of sequences that were recorded with a VGA (640x480) event camera (Samsung DVS Gen3) and a conventional RGB camera (Huawei P20 Pro) placed on the windshield of a car driving through Zurich.
1 paper · 0 benchmarks
Drone vs Bird (Drone vs Bird Detection Challenge)
For the Drone-vs-Bird Detection Challenge 2021, 77 different video sequences have been made available as training data.
1 paper · 1 benchmark
Dubbing Test Set consists of two subsets extracted from the En→De test set of COVOST-2, a large-scale multilingual speech translation corpus based on Common Voice.
1 paper · 0 benchmarks
We study dynamic appearance models of both relightable (BRDF) and non-relightable (RGB).
1 paper · 0 benchmarks
EDUVSUM (Educational Video Summarization)
EDUVSUM contains educational videos with subtitles from three popular e-learning platforms: Edx,YouTube, and TIB AV-Portal that cover the following topics: crash course on history of science and engineering, computer science, python and…
1 paper · 0 benchmarks
EE3P: Event-based Estimation of Periodic Phenomena Properties (Dataset) Kolář, J., Špetlík, R., Matas, J.
1 paper · 0 benchmarks
EGO-CH (EGOcentric-Cultural Heritage)
EGO-CH is a dataset of egocentric videos for visitors’ behavior understanding.
1 paper · 0 benchmarks
EGO-CH-Gaze (Learning to Detect Attended Objects in Cultural Sites with Gaze Signals and Weak Object Supervision)
To study the problem of weakly supervised attended object detection in cultural sites, we collected and labeled a dataset of egocentric images acquired from subjects visiting a cultural site.
1 paper · 0 benchmarks
ENF moving video (Electric Network Frequency Moving Video Dataset)
The ENF moving video dataset, which is a subset of the dataset used in Temporal Localization of Non-Static Digital Videos Using the Electrical Network Frequency , consists of video recording without the audio channel coupled with the…
1 paper · 0 benchmarks
EgoMon (Egomon Gaze & Video dataset)
EgoMon Gaze & Video Dataset is an Egocentric (first person) Dataset that consists of 7 videos of 30 minutes, more or less, each one of them.
1 paper · 0 benchmarks
The "Microbundle Time-lapse Dataset" contains 24 experimental time-lapse images of cardiac microbundles using three distinct types of experimental testbed of beating lab grown hiPSC-based cardiac microbundles.
1 paper · 0 benchmarks
The Extended UCF Crime extends the UCF Crime data set that consists of 13 anomaly classes.
1 paper · 0 benchmarks
The proposed Extended-YouTube Faces (E-YTF) is an extension of the famous YouTube Faces (YTF) dataset and is specifically designed to further push the challenges of face recognition by addressing the problem of open-set face identification…
1 paper · 0 benchmarks
The dataset X of this work is an extension of the heartSeg dataset.
1 paper · 1 benchmark
The EyeInfo Dataset is an open-source eye-tracking dataset created by Fabricio Batista Narcizo, a research scientist at the IT University of Copenhagen (ITU) and GN Audio A/S (Jabra), Denmark.
1 paper · 0 benchmarks
A large-scale isolated Indian sign language dataset.
1 paper · 1 benchmark
FISHTRAC (Nguyen Minh Khiem)
A dataset of real-world underwater videos annotated with multi-object tracking labels.
1 paper · 0 benchmarks
Fingertip Video Dataset of HB Estimation (Fingertip Video Dataset for Non-Invasive Diagnosis of Anemia)
This dataset comprises 1-minute fingertip video recordings collected from 150 anemic patients, ranging from 6 months to 32 years of age, with hemoglobin levels between 4.3 gm/dL and 12.4 gm/dL.
1 paper · 0 benchmarks
The Florentine dataset is a dataset of facial gestures which contains facial clips from 160 subjects (both male and female), where gestures were artificially generated according to a specific request, or genuinely given due to a shown…
1 paper · 0 benchmarks
FreeMan is the first large-scale multi-view human motion dataset under real scenarios.
1 paper · 0 benchmarks
The released GIF Reply dataset contains 1,562,701 real text-GIF conversation turns on Twitter.
1 paper · 1 benchmark
HA-ViD (HA-ViD: A Human Assembly Video Dataset)
Understanding comprehensive assembly knowledge from videos is critical for futuristic ultra-intelligent industry.
1 paper · 0 benchmarks
HT1080WT cells - 3D collagen type I matrices (HT1080WT cells embedded in 3D collagen type I matrices - manual annotations for cell instance segmentation and tracking)
Human fibrosarcoma HT1080WT (ATCC) cells at low cell densities embedded in 3D collagen type I matrices [1].
1 paper · 0 benchmarks
HYouTube is a video for Video harmonization, which aims to adjust the foreground of a composite video to make it compatible with the background.
1 paper · 0 benchmarks
Hawk Annotation Dataset includes language descriptions specifically for anomaly scenes in seven existing video anomaly datasets.
1 paper · 0 benchmarks
IAW Dataset (Ikea Assembly In The Wild Dataset)
The IAW dataset contains 420 Ikea furniture pieces from 14 common categories e.g.
1 paper · 0 benchmarks
Intelligent vehicle systems require a deep understanding of the interplay between road conditions, surrounding entities, and the ego vehicle's driving behavior for explainable driving decision-making and safe and efficient navigation.
1 paper · 0 benchmarks
IISc VINE (Indian Institute of Science VIdeo Naturalness Evaluation)
Indian Institute of Science VIdeo Naturalness Evaluation (IISc VINE) is a database consisting of 300 videos, obtained by applying different prediction models on different datasets, and accompanying human opinion scores.
1 paper · 0 benchmarks
INDRA (INdian Dataset for RoAd crossing)
INDRA is a dataset capturing videos of Indian roads from the pedestrian point-of-view.
1 paper · 0 benchmarks
InHARD (Industrial Human Action Recognition Dataset in the Context of Industrial Collaborative Robotics)
We introduce a RGB+S dataset named “Industrial Human Action Recognition Dataset” (InHARD) from a real-world setting for industrial human action recognition with over 2 million frames, collected from 16 distinct subjects.
1 paper · 0 benchmarks
Please find more details of this dataset at https://alex-xun-xu.github.io/ProjectPage/CVPR18/index.html 3D motion segmentation has been the key problem in computer vision research due to the application in structure from motion and…
1 paper · 1 benchmark
Kinetics-GEB+ (Generic Event Boundary Captioning, Grounding and Retrieval) is a dataset that consists of over 170k boundaries associated with captions describing status changes in the generic events in 12K videos.
1 paper · 3 benchmarks
A multimodal LIBRAS-UFOP Brazilian sign language dataset of minimal pairs using a microsoft Kinect senso.
1 paper · 1 benchmark
The LIRIS human activities dataset contains (gray/rgb/depth) videos showing people performing various activities taken from daily life (discussing, telphone calls, giving an item etc.).
1 paper · 0 benchmarks
LSA-T (Lengua de Señas Argentina - Traducción)
LSA-T is the first continuous Argentinian Sign Language (LSA) dataset.
1 paper · 1 benchmark
LSDBench (Long-video Sampling Dilemma Benchmark)
A benchmark that focuses on the sampling dilemma in long-video tasks.
1 paper · 0 benchmarks
LSFB Datasets (French Belgian Sign Language Datasets)
Sign Language Datasets for French Belgian Sign Language This dataset is built upon the work of Belgian linguists from the University of Namur.
1 paper · 0 benchmarks
The Large Scale Movie Description Challenge (LSMDC) - Context is an augmented version of the original LSMDC dataset with movie scripts as contextual text.
1 paper · 0 benchmarks
LTFT (Long-Term Face Tracking)
Dataset originally conceived for multi-face tracking/detection for highly crowded scenarios.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.