Home › Datasets › modality › Videos
Videos datasets
archive 2025-07-28
1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 20 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Videos datasets 913–960 of 1,014
TinyVIRAT-v2 is a benchmark dataset for recognizing real-world low-resolution activities present in videos.
1 paper · 0 benchmarks
TrUMAn (Trope Understanding in Movies and Animations)
Trope Understanding in Movies and Animations (TrUMAn) is a dataset intending to evaluate and develop learning systems beyond visual signals.
1 paper · 0 benchmarks
Trailers12k is a movie trailer dataset comprised of 12,000 titles associated to ten genres.
1 paper · 0 benchmarks
The VIriors Action Recognition Challenge uses a subset of the UCF101 action recognition dataset: Train set: ~4.8K clips.
1 paper · 0 benchmarks
Existing benchmark datasets in real-world distribution shifts are generally synthetically generated via augmentations to simulate real-world shifts such as weather and camera rotation.
1 paper · 0 benchmarks
UCF50 is an action recognition data set with 50 action categories, consisting of realistic videos taken from youtube.
1 paper · 0 benchmarks
Contains three difficult real-world scenarios: uncontrolled videos taken by UAVs and manned gliders, as well as controlled videos taken on the ground.
1 paper · 0 benchmarks
In this dataset UR5 robot used 6 tools: metal-scissor, metal-whisk, plastic-knife, plastic-spoon, wooden-chopstick, and wooden-fork to perform 6 behaviors: look, stirring-slow, stirring-fast, stirring-twist, whisk, and poke.
1 paper · 0 benchmarks
USC-GRAD-STDdb comprises 115 video segments containing more than 25,000 annotated frames of HD 720p resolution (≈1280x720) with small objects of interest from 16 (≈4x4) to 256 (≈16x16) as pixel area.
1 paper · 1 benchmark
We present UniTalk, a novel dataset specifically designed for the task of active speaker detection, emphasizing challenging scenarios to enhance model generalization.
1 paper · 0 benchmarks
A main goal of the Urban Soundscapes of the World project is to create a reference database of examples of urban acoustic environments, consisting of high-quality immersive audiovisual recordings (360-degree video and spatial audio), in…
1 paper · 0 benchmarks
V3C1 (the Vimeo Creative Commons Collection 1)
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, and will serve as evaluation basis for the Video Browser Showdown 2019-2021 and TREC Video Retrieval…
1 paper · 0 benchmarks
VCAS-Motion (Video Class Agnostic Segmentation Benchmark)
Video class agnostic segmentation (VCAS) is the task of segmenting objects without regards to its semantics combining appearance, motion and geometry from monocular video sequences.
1 paper · 0 benchmarks
VCG+112K (Video Instruction Dataset 112K)
Video-ChatGPT introduces the VideoInstruct100K dataset, which employs a semi-automatic annotation pipeline to generate 75K instruction-tuning QA pairs.
1 paper · 0 benchmarks
VETRA is a dataset for vehicle tracking in aerial image sequences and presents unique challenges such as low frame rates, small and fast-moving objects, as well as high camera movement.
1 paper · 0 benchmarks
VFD-2000 is a video fight detection dataset containing more than 2000 videos.
1 paper · 0 benchmarks
VILT (Video Instructions Linking for Complex Tasks)
VILT is a new benchmark collection of tasks and multimodal video content.
1 paper · 0 benchmarks
VISEM-Tracking is a dataset consisting of 20 video recordings of 30s of spermatozoa with manually annotated bounding-box coordinates and a set of sperm characteristics analyzed by experts in the domain.
1 paper · 0 benchmarks
VISOR is a dataset of pixel annotations and a benchmark suite for segmenting hands and active objects in egocentric video.
1 paper · 0 benchmarks
VOT2015 (Visual Object Tracking Challenge 2015)
VOT2015 is a visual object tracking dataset.
1 paper · 0 benchmarks
The benchmark for VPData, the largest video inpainting dataset, which comprises over 390K clips (> 866.7 hours) and features precise masks and detailed video captions.
1 paper · 0 benchmarks
The largest video inpainting dataset comprises over 390K clips (> 866.7 hours), featuring precise masks and detailed video captions.
1 paper · 0 benchmarks
VR-Folding contains garment meshes of 4 categories from CLOTH3D dataset, namely Shirt, Pants, Top and Skirt.
1 paper · 0 benchmarks
VSTaR-1M is a 1M instruction tuning dataset, created using Video-STaR, with the source datasets: Kinetics700 STAR-benchmark FineDiving The videos for VSTaR-1M can be found in the links above.
1 paper · 0 benchmarks
100 videos with varying danger levels (on a scale of 0-10) and different scenarios, annotated by 18 human annotators using our annotation pipeline to represent human perception and respective Vision Language model summaries for each of the…
1 paper · 0 benchmarks
We propose the first standardized benchmark in multimodal continual learning for video data, defining protocols for training and metrics for evaluation.
1 paper · 0 benchmarks
Introduction This dataset was gathered during the Vid2RealHRI study of humans’ perception of robots' intelligence in the context of an incidental Human-Robot encounter.
1 paper · 0 benchmarks
VidHarm is a professionally annotated dataset for detection of harmful content in video.
1 paper · 0 benchmarks
Vident-lab is a dataset of dental videos with multi-task labels to facilitate further research in relevant video processing applications.
1 paper · 0 benchmarks
Dataset Introduction This dataset leverages VideoDB's Public Collection to offer a diverse range of videos featuring text-containing scenes.
1 paper · 1 benchmark
Virtual-PedCross-4667 is a dataset for pedestrian crossing prediction.
1 paper · 0 benchmarks
This data contains about 2500 trajectories (with images and actions) of a Sawyer robot interacting with various objects.
1 paper · 0 benchmarks
W-Oops consists of 2,100 unintentional human action videos, with 44 goal-directed and 30 unintentional video-level activity labels collected through human annotations.
1 paper · 0 benchmarks
The Watch Your Mouth dataset is a custom silent speech dataset consisting of depth-only recordings of users silently mouthing full English sentences, captured using consumer-grade depth cameras such as the iPhone TrueDepth sensor.
1 paper · 0 benchmarks
A dataset automatically generated using question generation neural models and alt-text video captions from the WebVid dataset, with 3M video-question-answer triplets.
1 paper · 0 benchmarks
The dataset is a private dataset collected for automatic analysis of psychological distress.
1 paper · 1 benchmark
Werewolf Among Us is a dataset multimodal dataset for modeling persuasion behaviors.
1 paper · 0 benchmarks
WhenAct (Temporal Human Action Localization in Lifestyle Vlogs)
We consider the task of temporal human action localization in lifestyle vlogs.
1 paper · 0 benchmarks
WhyAct is a dataset for identifying human action reasons in online videos, consisting of 1,077 visual actions manually annotated with their reasons.
1 paper · 0 benchmarks
WiTA (Writing in The Air)
WiTA (Writing in The Air) is a dataset for the challenging writing in the air (WiTA) task -- an elaborate task bridging vision and NLP.
1 paper · 0 benchmarks
First of its kind paired win-fail action understanding dataset with samples from the following domains: “General Stunts,” “Internet Wins-Fails,” “Trick Shots,” & “Party Games.” The task is to identify successful and failed attempts at…
1 paper · 2 benchmarks
YouwikiHow is a dataset for Weakly-Supervised temporal Article Grounding (WSAG).
1 paper · 0 benchmarks
The first and the one open dataset for Russian finger- spelling, contained 1,593 annotated phrases and over 37 thousand HD+ videos.
1 paper · 1 benchmark
The i3-video dataset contains "is-it-instructional" annotations for 6.4k videos from Youtube-8M.
1 paper · 0 benchmarks
Please see our website and code repository for detailed description.
1 paper · 0 benchmarks
the dataset is a monkey doo doo dataset
1 paper · 0 benchmarks
Description: 895 Fire Videos Data,the total duration of videos is 27 hours 6 minutes 48.58 seconds.
0 papers · 0 benchmarks
ABODA (Abandoned Object Dataset)
ABandoned Objects DAtaset (ABODA) is a new public dataset for abandoned object detection.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.