Home › Datasets › modality › Videos

Videos datasets

archive 2025-07-28

1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 16 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Videos datasets 721–768 of 1,014

This dataset is meant to be used to develop models for next-day fire hazard forecasting in Greece.
1 paper · 0 benchmarks
Video samples recorded in the field using the Azure Kinect DK.
1 paper · 0 benchmarks
AMT Objects is a large dataset of object centric videos suitable for training and benchmarking models for generating 3D models of objects from a small number of photos of the objects.
1 paper · 0 benchmarks
ARKitTrack is a new RGB-D tracking dataset for both static and dynamic scenes captured by consumer-grade LiDAR scanners equipped on Apple's iPhone and iPad.
1 paper · 0 benchmarks
ASLLVD (American Sign Language Lexicon Video Dataset)
Extremely important: The ASLLVD video data are subject to Terms of Use: http://www.bu.edu/asllrp/signbank-terms.pdf.
1 paper · 0 benchmarks
A dataset for audio-visual event classification and localization in the context of office environments.
1 paper · 0 benchmarks
AVS Benchmark (Audio-Visual Synchrony Benchmark)
Provided in the linked paper.
1 paper · 0 benchmarks
AViMoS (Audio-Visual Mouse Saliency)
A novel audio-visual mouse saliency (AViMoS) dataset with the following key-features: Diverse content: movie, sports, live, vertical videos, etc.; Large scale: 1500 videos with mean 19s duration; High resolution: all streams are FullHD;…
1 paper · 0 benchmarks
Accidental Turntables contains a challenging set of 41,212 images of cars in cluttered backgrounds, motion blur and illumination changes that serves as a benchmark for 3D pose estimation.
1 paper · 0 benchmarks
Acticipate is a publicly available dataset with recordings of human body-motion and eye-gaze, acquired in an experimental scenario with an actor interacting with three subjects.
1 paper · 0 benchmarks
ActioNet is a video task-based dataset collected in a synthetic 3D environment.
1 paper · 0 benchmarks
Consists of 10,000+ video-sentence pairs with each accompanied by an annotated sentence specified video thumbnail.
1 paper · 0 benchmarks
AerialMPT is a dataset for pedestrian tracking in aerial image sequences and presents real-world challenges for MOT algorithms such as low frame rate, small moving objects, and complex backgrounds.
1 paper · 0 benchmarks
This large collection of over 161,000 video-label pairs of video clips, shows humans drawing letters and digits in the air, and is used to evaluate a model’s ability to classify articulated motions correctly.
1 paper · 0 benchmarks
Contains a large number of online videos and subtitles.
1 paper · 0 benchmarks
A small-scale, real-world Project Aria dataset with high quality static 3D oriented bounding boxs annotations.
1 paper · 1 benchmark
Aria Scenes is a benchmark dataset for future research on photorealistic reconstruction.
1 paper · 0 benchmarks
BAH (Behavioural Ambivalence/Hesitancy)
Recognizing complex emotions linked to ambivalence and hesitancy (A/H) can play a critical role in the personalization and effectiveness of digital behaviour change interventions.
1 paper · 0 benchmarks
BEAR (Benchmark on video Action Recognition)
BEAR (Benchmark on video Action Recognition) is a collection of 18 video datasets grouped into 5 categories (anomaly, gesture, daily, sports, and instructional), which covers a diverse set of real-world applications.
1 paper · 0 benchmarks
In this dataset two robots, Baxter and UR5, perform 8 behaviors (look, grasp, pick, hold, shake, lower, drop, and push) on 95 objects that vary by 5 color (blue, green, red, white, and yellow), 6 contents (wooden button, plastic dices,…
1 paper · 0 benchmarks
A dataset for flying honeybee detection introduced in "A Method for Detection of Small Moving Objects in UAV Videos".
1 paper · 1 benchmark
BioDrone is the first bionic drone-based single object tracking benchmark, it features videos captured from a flapping-wing UAV system with a major camera shake due to its aerodynamics.
1 paper · 0 benchmarks
Bukva (Bukva: Russian Sign Language Alphabet)
We introduce a video dataset Bukva for Russian Dactyl Recognition task.
1 paper · 1 benchmark
The feature files are named with the youtube IDs.
1 paper · 0 benchmarks
CANDOR Corpus (CANDOR = Conversation: A Naturalistic Dataset of Online Recordings)
The CANDOR corpus is a large, novel, multimodal corpus of 1,656 recorded conversations in spoken English.
1 paper · 0 benchmarks
CAP-DATA is a large-scale benchmark consisting of 11,727 in-the-wild accident videos with over 2.19 million frames together with labeled fact-effect-reason-introspection description and temporal accident frame label.
1 paper · 0 benchmarks
A new large-scale, in-thewild Mandarin dataset, CAS-VSR-S101 with 101.1 hours of data.
1 paper · 3 benchmarks
CASR (Cyclist Arm Signal Recognition)
CASR is a dataset for cyclist arm signal recognition in videos.
1 paper · 0 benchmarks
CLAD (Complex and Long Activities Dataset)
CLAD (Compled and Long Activities Dataset) is an activity dataset which exhibits real-life and diverse scenarios of complex, temporally-extended human activities and actions.
1 paper · 0 benchmarks
CN-Celeb-AV is a multi-genre AVPR dataset collected 'in the wild'.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
CRIM13 (Caltech Resident-Intruder Mouse 13)
The Caltech Resident-Intruder Mouse dataset (CRIM13) consists of 237x2 videos (recorded with synchronized top and side view) of pairs of mice engaging in social behavior, catalogued into thirteen different actions.
1 paper · 0 benchmarks
In this dataset an uppertorso humanoid robot with 7-DOF arm explored 100 different objects belonging to 20 different categories using 10 behaviors: Look, Crush, Grasp, Hold, Lift, Drop, Poke, Push, Shake and Tap.
1 paper · 0 benchmarks
This dataset comprises video files (converted into tif format) that depict glomerular activation in mice.
1 paper · 0 benchmarks
57 stock videos from Pexels, predominantly covering road scenes which involve minimal distortion.
1 paper · 0 benchmarks
ChinaOpen is a new video dataset targeted at open-world multimodal learning, with raw data gathered from Bilibili, a popular Chinese video-sharing website.
1 paper · 1 benchmark
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
1 paper · 0 benchmarks
ConSLAM (Construction Dataset for SLAM)
ConSLAM is a real-world dataset collected periodically on a construction site to measure the accuracy of mobile scanners' SLAM algorithms.
1 paper · 0 benchmarks
This is a video and image segmentation dataset for human head and shoulders, relevant for creating elegant media for videoconferencing and virtual reality applications.
1 paper · 0 benchmarks
Description - Repository: Code, Page, Data - Paper: arxiv.org/abs/2411.17440 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a star and citation.
1 paper · 0 benchmarks
This spatio-temporal actions dataset for video understanding consists of 4 parts: original videos, cropped videos, video frames, and annotation files.
1 paper · 0 benchmarks
D-OCC (Dynamic-OneCommon Corpus)
D-OCC is a large-scale dataset of 5,617 dialogues to enable fine-grained evaluation and analysis of various dialogue systems.
1 paper · 0 benchmarks
A large-scale comprehensive collection of dashcam videos collected by vehicles on DiDi's platform.
1 paper · 0 benchmarks
DADE (Driving Agents in Dynamic Environments)
The DADE dataset, short for Driving Agents in Dynamic Environments, is a synthetic dataset designed for the training and evaluation of methods for the task of semantic segmentation in the context of autonomous driving agents navigating…
1 paper · 0 benchmarks
DARai (Daily Activity Recordings for AI and ML applications)
Daily Activity Recordings for Artificial Intelligence (DARai, pronounced "Dahr-ree") is a multimodal, hierarchically annotated dataset constructed to understand human activities in real-world settings.
1 paper · 0 benchmarks
DAVIDE ('Depth-Aware VIdeo DEblurring')
The DAVIDE dataset consists of synchronized blurred, depth, and sharp videos.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.