Home › Datasets › modality › Videos

Videos datasets

archive 2025-07-28

1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 14 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Videos datasets 625–672 of 1,014

V-HICO is a dataset for human-object interaction in videos.
3 papers · 0 benchmarks
VATEX Adverbs is a subset from VATEX with extracted verb-adverb annotations.
3 papers · 2 benchmarks
Vehicle-Rear is a novel dataset for vehicle identification that contains more than three hours of high-resolution videos, with accurate information about the make, model, color and year of nearly 3,000 vehicles, in addition to the position…
3 papers · 0 benchmarks
VideoMatting108 is a large-scale video matting and trimap generation dataset with 80 training and 28 validation foreground video clips with ground-truth alpha mattes.
3 papers · 0 benchmarks
The WebVid-CoVR dataset is a collection of video-text-video triplets that can be used for the task of composed video retrieval (CoVR).
3 papers · 1 benchmark
YTD-18M is a large-scale corpus of 18M video-based dialogues, constructed from web videos: crucial to the data collection pipeline is a pretrained language model that converts error-prone automatic transcripts to a cleaner dialogue format…
3 papers · 0 benchmarks
3DYoga90 (3DYoga90: A Hierarchical Video Dataset for Yoga Pose Understanding)
3DYoga90 is organized within a three-level label hierarchy.
2 papers · 0 benchmarks
A multilingual, multimodal and multi-aspect, expertly-annotated dataset of diverse short videos extracted from short-video social media platform - Moj.
2 papers · 0 benchmarks
ASD (Annotated Semantic Dataset)
The Annotated Semantic Dataset is composed of $11$ videos, divided in $3$ activity categories: Biking; Driving and Walking, according to their amount of semantic information.
2 papers · 0 benchmarks
AVCAffe (A Large Scale Audio-Visual Dataset of Cognitive Load and Affect for Remote Work)
We introduce AVCAffe, the first Audio-Visual dataset consisting of Cognitive load and Affect attributes.
2 papers · 0 benchmarks
AVMIT (Audiovisual Moments in Time)
Audiovisual Moments in Time (AVMIT) is a large-scale dataset of audiovisual action events.
2 papers · 0 benchmarks
AVSync15 is a high-quality synchronized audio-video dataset curated from VGGSound.
2 papers · 0 benchmarks
ActionBench contains two carefully designed probing tasks: Action Antonym and Video Reversal, which targets multimodal alignment capabilities and temporal understanding skills of the model, respectively.
2 papers · 0 benchmarks
The Argoverse 2 Map Change Dataset is a collection of 1,000 scenarios with ring camera imagery, lidar, and HD maps.
2 papers · 0 benchmarks
A real-world dataset, with hyper-accurate digital counterpart & comprehensive ground-truth annotation.
2 papers · 1 benchmark
The availability of high-quality datasets play a crucial role in advancing research and development especially, for safety critical and autonomous systems.
2 papers · 0 benchmarks
BioVid (BioVid Heat Pain Database)
To advance methods for pain assessment, in particular automatic assessment methods, the BioVid Heat Pain Database was collected in a collaboration of the Neuro-Information Technology group of the University of Magdeburg and the Medical…
2 papers · 0 benchmarks
The breast lesion detection in ultrasound videos dataset uses a clip-level and video-level feature aggregated network (CVA-Net) and consists of 188 ultrasound videos, of which 113 are labeled malignant and 75 benign.
2 papers · 0 benchmarks
CAP (Consented Activities of People)
The Consented Activities of People (CAP) dataset is a fine grained activity dataset for visual AI research curated using the Visym Collector platform.
2 papers · 0 benchmarks
Our CCTV-Pipe dataset consists of 16 defect categories including structural and functional defects in the pipe.
2 papers · 0 benchmarks
CRIPP-VQA (Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering)
CRIPP-VQA is a video question answering dataset for reasoning about the implicit physical properties of objects in a scene.
2 papers · 0 benchmarks
CVB (Video Dataset of Cattle Visual Behaviors)
Existing image/video datasets for cattle behavior recognition are mostly small, lack well-defined labels, or are collected in unrealistic controlled environments.
2 papers · 0 benchmarks
CholecT40 (Cholecystectomy Action Triplet)
CholecT40 is the first endoscopic dataset introduced to enable research on fine-grained action recognition in laparoscopic surgery.
2 papers · 1 benchmark
CholecTrack20 (Multi-Perspective Multi-Class Multi-Object Tracking Dataset For Surgical Tools)
CholecTrack20 is a surgical video dataset focusing on laparoscopic cholecystectomy and designed for surgical tool tracking, featuring 20 annotated videos.
2 papers · 0 benchmarks
Description - Venue: NeurIPS 2024 D&B Spotlight - Repository: Code, Page, Data - Paper: arxiv.org/abs/2406.18522 - Point of Contact: Shenghai Yuan Citation If you find our paper and code useful in your research, please consider giving a…
2 papers · 0 benchmarks
Cinescale (CineScale: A dataset of cinematic shot scale in movies)
We provide a database containing shot scale annotations (i.e., the apparent distance of the camera from the subject of a filmed scene) for more than 792,000 image frames.
2 papers · 0 benchmarks
DL3DV-10K is a dataset of real-world videos with scene annotations and camera parameters.
2 papers · 0 benchmarks
DeePhy is a novel DeepFake Phylogeny dataset consisting of 5040 DeepFake videos generated using three different generation techniques.
2 papers · 0 benchmarks
Drone-Action (Drone-Action: An Outdoor Recorded Drone Video Dataset for Action Recognition)
Website: https://asankagp.github.io/droneaction/
2 papers · 1 benchmark
DropletVideo is a project exploring high-order spatio-temporal consistency in image-to-video generation.
2 papers · 0 benchmarks
Estimating camera motion in deformable scenes poses a complex and open research challenge.
2 papers · 1 benchmark
Dynamic OLAT Dataset (ShanghaiTech MARS Dynamic OLAT Dataset)
To provide ground truth supervision for video consistency modeling, we build up a high-quality dynamic OLAT dataset.
2 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
2 papers · 1 benchmark
Contains annotations of human activity with different sub-actions, e.g., activity Ping-Pong with four sub-actions which are pickup-ball, hit, bounce-ball and serve.
2 papers · 0 benchmarks
ERATO is a large-scale multi-modal dataset for Pairwise Emotional Relationship Recognition (PERR).
2 papers · 0 benchmarks
The fetoscopy placenta dataset is associated with our MICCAI2020 publication titled “Deep Placental Vessel Segmentation for Fetoscopic Mosaicking”.
2 papers · 0 benchmarks
GAFA (Gaze from afar dataset)
We introduce a new dataset of annotated surveillance videos of freely moving people taken from a distance in both indoor and outdoor scenes.
2 papers · 0 benchmarks
We release the dataset for non-commercial research.
2 papers · 0 benchmarks
HAMMER dataset contains 13 Scenes.
2 papers · 0 benchmarks
This dataset contains panoramic video captured from a helmet-mounted camera while riding a bike through suburban Northern Virginia.
2 papers · 0 benchmarks
Whereas the action recognition community has focused mostly on detecting simple actions like clapping, walking or jogging, the detection of fights or in general aggressive behaviors has been comparatively less studied.
2 papers · 1 benchmark
I2-2000FPS is the first high-speed video dataset offering an unprecedented temporal resolution of 2000 frames per second (fps).
2 papers · 0 benchmarks
IACC.3 (Internet Archive videos (IACC.3) under Creative Commons licenses.)
The IACC.3 dataset is approximately 4600 Internet Archive videos (144 GB, 600 h) with Creative Commons licenses in MPEG-4/H.264 format with duration ranging from 6.5 min to 9.5 min and a mean duration of almost 7.8 min.
2 papers · 0 benchmarks
Sign languages are the primary means of communication for a large number of people worldwide.
2 papers · 0 benchmarks
ITB (Informative Tracking Benchmark)
Informative Tracking Benchmark (ITB) is a small and informative tracking benchmark with 7% out of 1.2 M frames of existing and newly collected datasets, which enables efficient evaluation while ensuring effectiveness.
2 papers · 1 benchmark
JRDB-Pose is a large-scale dataset and benchmark for multi-person pose estimation and tracking using videos captured from a social navigation robot.
2 papers · 0 benchmarks
This is a subset of Kinetics-400, introduced in Look, Listen and Learn by Relja Arandjelovic and Andrew Zisserman.
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.