Home › Datasets › modality › Videos
Videos datasets
archive 2025-07-28
1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 21 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Videos datasets 961–1008 of 1,014
BAVL (Blind Audio-Visual Localization (BAVL))
Blind Audio-Visual Localization (BAVL) Dataset consists of 20 audio-visual recordings of sound sources, which could be talking faces or music instruments.
0 papers · 0 benchmarks
BMS-26 (Berkeley Motion Segmentation)
The Berkeley Motion Segmentation Dataset (BMS-26) is a dataset for motion segmentation, which consists of 26 video sequences with pixel-accurate segmentation annotation of moving objects.
0 papers · 0 benchmarks
The CTV-Dataset (CTV stands for Cyclist Top-View) is a trajectories dataset for cyclist behaviour in mixed-traffic environments (aka.
0 papers · 0 benchmarks
CUHK Square data set is for transfer learning research on adapting generic pedestrian detectors.
0 papers · 0 benchmarks
This dataset focuses only on the robbery category, presenting a new weakly labelled dataset that contains 486 new real–world robbery surveillance videos acquired from public sources.
0 papers · 0 benchmarks
The Couples Therapy corpus contains audio, video recordings and manual transcriptions of conversations between 134 real-life couples attending marital therapy.
0 papers · 0 benchmarks
Crowd 11 (A Dataset for Fine Grained Crowd Behaviour Analysis)
This dataset defines a total of 11 crowd motion patterns and it is composed of over 6000 video sequences with an average length of 100 frames per sequence.
0 papers · 0 benchmarks
DAHLIA (DAily Human Life Activity)
DAHLIA dataset [1] is devoted to human activity recognition, which is a major issue for adapting smart-home services such as user assistance.
0 papers · 0 benchmarks
Dataset for the DREAMING - Diminished Reality for Emerging Applications in Medicine through Inpainting Challenge!
0 papers · 0 benchmarks
DUS (Daimler Urban Segmentation)
The Daimler Urban Segmentation Dataset is a dataset for semantic segmentation.
0 papers · 0 benchmarks
The DeepSpeak dataset contains over 43 hours of real and deepfake footage of people talking and gesturing in front of their webcams.
0 papers · 0 benchmarks
The DogCentric Activity dataset is composed of dog activity videos taken from a first-person animal viewpoint.
0 papers · 0 benchmarks
FPV-O is a multi-subject first-person vision dataset of office activities.
0 papers · 0 benchmarks
We introduce FortisAVQA, a dataset designed to assess the robustness of AVQA models.
0 papers · 0 benchmarks
Freiburg Block Tasks is a dataset for robot skill learning.
0 papers · 0 benchmarks
The Freiburg Poking dataset is a dataset for learning intuitive physics from physical interaction.
0 papers · 0 benchmarks
Gap Pattern Detection (Gap Pattern (Gap Up and Gap Down) Detection in Candlestick Trading Charts for Technical Analysis)
1.
0 papers · 0 benchmarks
GenAI-Bench benchmark consists of 1,600 challenging real-world text prompts sourced from professional designers.
0 papers · 0 benchmarks
The Grand central station dataset includes a video with 50,010 frames which is used for Scene Understanding and Crowd Analysis.
0 papers · 0 benchmarks
HEADSET (HEADSET: Human Emotion Awareness under Partial Occlusions Multimodal DataSET)
The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications.
0 papers · 0 benchmarks
The medaka (Oryzias latipes) and the zebrafish (Danio rerio) are used as a model organism for a variety of subjects in biomedical research.
0 papers · 0 benchmarks
IMO (Independently Moving Objects)
Dataset of annotated independently moving objects (IMO).
0 papers · 0 benchmarks
InfiniteRep is a synthetic, open-source dataset for fitness and physical therapy (PT) applications.
0 papers · 0 benchmarks
Infinity AI's Spills Basic Dataset is a synthetic, open-source dataset for safety applications.
0 papers · 0 benchmarks
To study kinship verification from gait, we collected the dataset KinGaitWild consisting of several videos from youtube.
0 papers · 0 benchmarks
L-SVD (Large-Scale Selfie Video Dataset (L-SVD): A Benchmark for Emotion Recognition)
Welcome to L-SVD L-SVD is an extensive and rigorously curated video dataset aimed at transforming the field of emotion recognition.
0 papers · 0 benchmarks
LASIESTA (Labeled and Annotated Sequences for Integral Evaluation of SegmenTation Algorithms) is a segmentation and detection dataset composed by many real indoor and outdoor sequences organized into categories, each of one covering a…
0 papers · 0 benchmarks
This is a dataset for vehicle detection.
0 papers · 0 benchmarks
MCCSD (Mandarin Chinese Cued Speech Dataset)
This MCCS dataset is the first large-scale Mandarin Chinese Cued Speech dataset.
0 papers · 0 benchmarks
The MOBIO database consists of bi-modal (audio and video) data taken from 152 people.
0 papers · 0 benchmarks
The MS-EVS Dataset is the first large-scale event-based dataset for face detection.
0 papers · 0 benchmarks
MoCap (CMU Graphics Lab Motion Capture Database)
Collection of various motion capture recordings (walking, dancing, sports, and others) performed by over 140 subjects.
0 papers · 0 benchmarks
The Mouse Embryo Tracking Database is a dataset for tracking mouse embryos.
0 papers · 0 benchmarks
A pedestrian dataset for Person Re-identification.
0 papers · 0 benchmarks
OCTCBVS is a benchmark dataset for testing and evaluating novel and state-of-the-art computer vision algorithms.
0 papers · 0 benchmarks
The OpenEQA dataset is a significant contribution in the field of Embodied Question Answering (EQA).
0 papers · 0 benchmarks
The PIROPO database (People in Indoor ROoms with Perspective and Omnidirectional cameras) comprises multiple sequences recorded in two different indoor rooms, using both omnidirectional and perspective cameras.
0 papers · 0 benchmarks
Plant Centroids is a dataset for stem emerging points (SEP) detection in RGB and NIR image data.
0 papers · 0 benchmarks
Laser powder bed fusion (LBPF) is the additive manufacturing (3D printing) process for metals.
0 papers · 0 benchmarks
SICS-155 (Phase Recognition in Small Incision Cataract Surgery Videos)
Cataract is the leading cause of blindness worldwide, most affecting life in low- and middle-income countries (LMICs).
0 papers · 0 benchmarks
STVD-FC is the largest public dataset on the political content analysis and fact-checking tasks.
0 papers · 0 benchmarks
It is released by the Shanghai Central Meteorological Observatory (SCMO) in 2020, records serval years of historical precipitation events in the Yangtze River delta area.
0 papers · 0 benchmarks
THVD (Talking Head Video Dataset)
About We provide a comprehensive talking-head video dataset with over 50,000 videos, totaling more than 500+ hours of footage and featuring 20,841 unique identities from around the world.
0 papers · 0 benchmarks
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips,…
0 papers · 0 benchmarks
The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection.
0 papers · 0 benchmarks
Toronto NeuroFace Dataset: A New Dataset for Facial Motion Analysis in Individuals with Neurological Disorders Toronto NeuroFace Dataset is a public dataset with videos of oro-facial gestures performed by individuals with oro-facial…
0 papers · 0 benchmarks
Human activity recognition and clinical biomechanics are challenging problems in physical telerehabilitation medicine.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.