Home › Datasets › modality › Videos
Videos datasets
archive 2025-07-28
1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 10 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Videos datasets 433–480 of 1,014
LDV (Large-scale Diverse Video)
LDV is a dataset for video enhancement.
7 papers · 0 benchmarks
LSVTD is a large scale video text dataset for promoting the video text spotting community, which contains 100 text videos from 22 different real-life scenarios.
7 papers · 0 benchmarks
This work proposes Long-RVOS, a large-scale benchmark for long-term video object segmentation.
7 papers · 1 benchmark
MSU BASED (MSU BASED Video Deblurring Dataset and Benchmark)
Qualitative dataset with real blurred videos, created by using beam-splitter setup in lab environment
7 papers · 1 benchmark
MonoPerfCap is a benchmark dataset for human 3D performance capture from monocular video input consisting of around 40k frames, which covers a variety of different scenarios.
7 papers · 0 benchmarks
NExT-QA is a VideoQA benchmark targeting the explanation of video contents.
7 papers · 1 benchmark
NTU RGB+D 2D is a curated version of NTU RGB+D often used for skeleton-based action prediction and synthesis.
7 papers · 1 benchmark
ORBIT is a real-world few-shot dataset and benchmark grounded in a real-world application of teachable object recognizers for people who are blind/low vision.
7 papers · 2 benchmarks
The Privacy Annotated HMDB51 (PA-HMDB51) dataset is a video-based dataset for evaluating pirvacy protection in visual action recognition algorithms.
7 papers · 0 benchmarks
POT-210 (Planar Object Tracking in the Wild: A Benchmark)
Planar object tracking is an actively studied problem in vision-based robotic applications.
7 papers · 0 benchmarks
SVD (Short Video Dataset)
SVD is a large-scale short video dataset, which contains over 500,000 short videos collected from http://www.douyin.com and over 30,000 labeled pairs of near-duplicate videos.
7 papers · 0 benchmarks
SynPick is a synthetic dataset for dynamic scene understanding in bin-picking scenarios.
7 papers · 1 benchmark
The SynthHands dataset is a dataset for hand pose estimation which consists of real captured hand motion retargeted to a virtual hand with natural backgrounds and interactions with different objects.
7 papers · 0 benchmarks
TAPE (resToration of digitized Analog videotaPEs)
A dataset of videos synthetically degraded with Adobe After Effects to exhibit artifacts resembling those of real-world analog videotapes.
7 papers · 1 benchmark
TREK-150 is a benchmark dataset for object tracking in First Person Vision (FPV) videos composed of 150 densely annotated video sequences.
7 papers · 0 benchmarks
The TUM Kitchen dataset is an action recognition dataset that contains 20 video sequences captured by 4 cameras with overlapping views.
7 papers · 0 benchmarks
UBI-Fights - Concerning a specific anomaly detection and still providing a wide diversity in fighting scenarios, the UBI-Fights dataset is a unique new large-scale dataset of 80 hours of video fully annotated at the frame level.
7 papers · 2 benchmarks
VPCD (Video Person-Clustering)
VPCD contains multi-modal annotations (face, body and voice) for all primary and secondary characters from a range of diverse TV-shows and movies.
7 papers · 1 benchmark
VidHOI is a video-based human-object interaction detection benchmark.
7 papers · 2 benchmarks
WALT (Watch and Learn TimeLapse Images)
We introduce a new dataset, Watch and Learn Time-lapse (WALT), consisting of multiple (4K and 1080p) cameras capturing urban environments over a year.
7 papers · 1 benchmark
YT-BB (YouTube-BoundingBoxes)
YouTube-BoundingBoxes (YT-BB) is a large-scale data set of video URLs with densely-sampled object bounding box annotations.
7 papers · 1 benchmark
iPhone dataset is a challenging benchmarks for dynamic reconstruction.
7 papers · 1 benchmark
300-VW (300 Videos in the Wild)
300 Videos in the Wild (300-VW) is a dataset for evaluating facial landmark tracking algorithms in the wild.
6 papers · 2 benchmarks
ARID is a dataset for action recognition in dark videos.
6 papers · 0 benchmarks
The Bimanual Actions Dataset is a collection of 540 RGB-D videos, showing subjects perform bimanual actions in a kitchen or workshop context.
6 papers · 0 benchmarks
DroneCrowd is a benchmark for object detection, tracking and counting algorithms in drone-captured videos.
6 papers · 0 benchmarks
DroneSURF (DroneSURF: Benchmark Dataset for Drone-based Face Recognition)
Drone Surveillance of Faces, is a large-scale drone dataset intended to facilitate research for face recognition using drones.
6 papers · 1 benchmark
This dataset contains around 10000 videos generated by various methods using the Prompt list.
6 papers · 1 benchmark
The German Lipreading dataset consists of 250,000 publicly available videos of the faces of speakers of the Hessian Parliament, which was processed for word-level lip reading using an automatic pipeline.
6 papers · 0 benchmarks
IndustReal (IndustReal Dataset of Egocentric Videos for Procedure Understanding)
IndustReal is an ego-centric, multi-modal dataset where 27 participants are challenged to perform assembly and maintenance procedures on a construction-toy car.
6 papers · 3 benchmarks
MMPTRACK (Multi-camera Multiple People Tracking Dataset)
Multi-camera Multiple People Tracking (MMPTRACK) dataset has about 9.6 hours of videos, with over half a million frame-wise annotations.
6 papers · 1 benchmark
A quantitative benchmark for developing and understanding video of fill-in-the-blank question-answering dataset with over 300,000 examples, based on descriptive video annotations for the visually impaired.
6 papers · 0 benchmarks
MovieShots is a dataset to facilitate the shot type analysis in videos.
6 papers · 0 benchmarks
Moviescope is a large-scale dataset of 5,000 movies with corresponding video trailers, posters, plots and metadata.
6 papers · 0 benchmarks
The ObjectFolder Real dataset contains multisensory data collected from 100 real-world household objects.
6 papers · 0 benchmarks
The dataset is split between train, test and val folders.
6 papers · 0 benchmarks
A dataset for text in driving videos.
6 papers · 0 benchmarks
TRECVID is a yearly set of competitions centered on video retrieval and indexing, hosting a variety of video data sets.
6 papers · 1 benchmark
V2C (Video-to-Commonsense)
6 papers · 0 benchmarks
VIPER is a benchmark suite for visual perception.
6 papers · 0 benchmarks
VideoCube is a high-quality and large-scale benchmark to create a challenging real-world experimental environment for Global Instance Tracking (GIT).
6 papers · 1 benchmark
VideoLT is a large-scale long-tailed video recognition dataset that contains 256,218 untrimmed videos, annotated into 1,004 classes with a long-tailed distribution.
6 papers · 0 benchmarks
WebLINX (Real-World Website Navigation with Multi-Turn)
WebLINX is a large-scale benchmark of 100K interactions across 2300 expert demonstrations of conversational web navigation.
6 papers · 1 benchmark
A new dataset with significant occlusions related to object manipulation.
6 papers · 0 benchmarks
Video object segmentation has been studied extensively in the past decade due to its importance in understanding video spatial-temporal structures as well as its value in industrial applications.
6 papers · 1 benchmark
iQIYI-VID dataset, which comprises video clips from iQIYI variety shows, films, and television dramas.
6 papers · 0 benchmarks
Acappella comprises around 46 hours of a cappella solo singing videos sourced from YouTbe, sampled across different singers and languages.
5 papers · 0 benchmarks
BRACE (The Breakdancing Competition Dataset for Dance Motion Synthesis)
BRACE is a dataset for audio-conditioned dance motion synthesis challenging common assumptions for this task: - strong music-dance correlation - controlled motion data - simple poses and movements To address these issues: - We focus on…
5 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.