Home › Datasets › modality › Videos

Videos datasets

archive 2025-07-28

1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 11 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Videos datasets 481–528 of 1,014

BTS3.1 (Expanding Accurate Person Recognition to New Altitudes and Ranges: The BRIAR Dataset)
Large, multimodal biometric dataset: It contains still images and videos of over 1,000 people captured at various ranges (up to 1,000 meters) and elevations (up to 400 meters) using a diverse set of cameras (commercial, military-grade,…
5 papers · 2 benchmarks
BUP20 (Sweet Pepper 2020 University of Bonn)
Video sequences from a glasshouse environment in Campus Kleinaltendorf(CKA), University of Bonn, captured by PATHoBot, a glasshouse monitoring robot.
5 papers · 0 benchmarks
ChronoMagic with 2265 metamorphic time-lapse videos, each accompanied by a detailed caption.
5 papers · 0 benchmarks
CoP3D is a collection of crowd-sourced videos showing around 4,200 distinct pets.
5 papers · 0 benchmarks
ComPhy (Compositional Physical Reasoning Dataset)
Compositional Physical Reasoning is a dataset for understanding object-centric and relational physics properties hidden from visual appearances.
5 papers · 0 benchmarks
D3D-HOI is a dataset of monocular videos with ground truth annotations of 3D object pose, shape and part motion during human-object interactions.
5 papers · 0 benchmarks
The Distress Analysis Interview Corpus/Wizard-of-Oz set (DAIC-WOZ) dataset [24, 25] comprises voice and text samples from 189 interviewed healthy and control persons and their PHQ-8 depression detection questionnaire.
5 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
5 papers · 0 benchmarks
We construct the ForgeryNet dataset, an extremely large face forgery dataset with unified annotations in image- and video-level data across four tasks: 1) Image Forgery Classification, including two-way (real / fake), three-way (real /…
5 papers · 1 benchmark
The first AI-generated video detection datasets.
5 papers · 0 benchmarks
GolfDB is a high-quality video dataset created for general recognition applications in the sport of golf, and specifically for the task of golf swing sequencing.
5 papers · 0 benchmarks
HARPER (Exploring 3D Human Pose Estimation and Forecasting from the Robot’s Perspective: The HARPER Dataset)
We introduce HARPER, a novel dataset for 3D body pose estimation and forecast in dyadic interactions between users and \spot, the quadruped robot manufactured by Boston Dynamics.
5 papers · 3 benchmarks
HEV-I (Honda Egocentric View-Intersection Dataset)
Honda Egocentric View-Intersection Dataset (HEV-I) is introduced to enable research on traffic participants interaction modelling, future object localization, as well as learning driver action in challenging driving scenarios.
5 papers · 1 benchmark
We introduce HourVideo, a benchmark dataset for hour-long video-language understanding.
5 papers · 0 benchmarks
A dataset of 69,270,581 video clip, question and answer triplets (v, q, a).
5 papers · 0 benchmarks
The ImplicitQA dataset was introduced in the paper ImplicitQA: Going beyond frames towards Implicit Video Reasoning.
5 papers · 1 benchmark
We collect, organize and open-source the large-scale multimodal instruction dataset, Infinity-MM, consisting of tens of millions of samples.
5 papers · 0 benchmarks
JRDB-Act is an extension of the JRDB dataset to create a large-scale multi-modal dataset for spatio-temporal action, social group and activity detection.
5 papers · 0 benchmarks
LIVE Livestream is a database for Video Quality Assessment (VQA), specifically designed for live streaming VQA research.
5 papers · 1 benchmark
LaRS (Lakes, Rivers and Seas Dataset)
LaRS is the largest and most diverse panoptic maritime obstacle detection dataset.
5 papers · 3 benchmarks
LoLi-Phone is a large-scale low-light image and video dataset for Low-light image enhancement (LLIE).
5 papers · 0 benchmarks
The MLB-YouTube dataset is a new, large-scale dataset consisting of 20 baseball games from the 2017 MLB post-season available on YouTube with over 42 hours of video footage.
5 papers · 0 benchmarks
A dataset which provides detailed annotations for activity recognition.
5 papers · 1 benchmark
This is a dataset for a shot boundary detection task.
5 papers · 1 benchmark
Understanding what makes a video memorable has a very broad range of current applications, e.g., education and learning, content retrieval and search, content summarization, storytelling, targeted advertising, content recommendation and…
5 papers · 0 benchmarks
Synthetically Generated Night-time Weather Degraded Database
5 papers · 1 benchmark
OAK (Objects Around Krishna)
OAK is a dataset for online continual object detection benchmark with an egocentric video dataset.
5 papers · 0 benchmarks
RF100 (Roboflow 100)
The evaluation of object detection models is usually performed by optimizing a single metric, e.g.
5 papers · 1 benchmark
RoCoG-v2 (Robot Control Gestures)
RoCoG-v2 (Robot Control Gestures) is a dataset intended to support the study of synthetic-to-real and ground-to-air video domain adaptation.
5 papers · 1 benchmark
Specially designed to evaluate active learning for video object detection in road scenes.
5 papers · 0 benchmarks
The Sims4Action Dataset: a videogame-based dataset for Synthetic→Real domain adaptation for human activity recognition.
5 papers · 0 benchmarks
This data collection consists of images acquired during chemoradiotherapy of 20 locally-advanced, non-small cell lung cancer patients.
5 papers · 0 benchmarks
Internet Archive videos (IACC.3) under Creative Commons licenses.
5 papers · 1 benchmark
Internet Archive videos (IACC.3) under Creative Commons licenses.
5 papers · 1 benchmark
Internet Archive videos (IACC.3) under Creative Commons licenses.
5 papers · 1 benchmark
A collection of 2511 recipes for zero-shot learning, recognition and anticipation.
5 papers · 0 benchmarks
Test-of-Time (Test of Time Synthetic Video Dataset)
The goal of this dataset is to probe video-language models for understanding of simple temporal relations like "before" and "after".
5 papers · 1 benchmark
TutorialVQA is a new type of dataset used to find answer spans in tutorial videos.
5 papers · 0 benchmarks
The UCLA Aerial Event Dataest has been captured by a low-cost hex-rotor with a GoPro camera, which is able to eliminate the high frequency vibration of the camera and hold in air autonomously through a GPS and a barometer.
5 papers · 0 benchmarks
UESTC RGB-D (UESTC RGB-D Varying-view action database)
UESTC RGB-D Varying-view action database contains 40 categories of aerobic exercise.
5 papers · 1 benchmark
VCSL (Video Copy Segment Localization)
VCSL (Video Copy Segment Localization) is a new comprehensive segment-level annotated video copy dataset.
5 papers · 0 benchmarks
The dataset uses VGG-Sound which consists of 10s clips collected from YouTube for 309 sound classes.
5 papers · 0 benchmarks
A large-scale multi-modal dataset to facilitate research and studies that concentrate on vision-wireless systems.
5 papers · 1 benchmark
VidChapters-7M is a dataset of 817K user-chaptered videos including 7M chapters in total.
5 papers · 4 benchmarks
VidOR (Video Object Relation) dataset contains 10,000 videos (98.6 hours) from YFCC100M collection together with a large amount of fine-grained annotations for relation understanding.
5 papers · 1 benchmark
Due to the lack of training data for video waterdrop removal, we propose a large-scale synthetic dataset with simulated waterdrops in complex driving scenes on rainy days.
5 papers · 1 benchmark
WEAR (WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition)
WEAR is an outdoor sports dataset for both vision- and inertial-based human activity recognition (HAR).
5 papers · 0 benchmarks
WWW Crowd provides 10,000 videos with over 8 million frames from 8,257 diverse scenes, therefore offering a comprehensive dataset for the area of crowd understanding.
5 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.