Home › Datasets › modality › Videos
Videos datasets
archive 2025-07-28
1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 7 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Videos datasets 289–336 of 1,014
The FIVR-200K dataset has been collected to simulate the problem of Fine-grained Incident Video Retrieval (FIVR).
16 papers · 1 benchmark
Jester Gesture Recognition dataset includes 148,092 labeled video clips of humans performing basic, pre-defined hand gestures in front of a laptop camera or webcam.
16 papers · 6 benchmarks
MeetingBank, a benchmark dataset created from the city councils of 6 major U.S.
16 papers · 1 benchmark
TSU (Toyota Smarthome Untrimmed)
Toyota Smarthome Untrimmed (TSU) is a dataset for activity detection in long untrimmed videos.
16 papers · 1 benchmark
TV show Caption is a large-scale multimodal captioning dataset, containing 261,490 caption descriptions paired with 108,965 short video moments.
16 papers · 1 benchmark
VideoLQ consists of videos downloaded from various video hosting sites such as Flickr and YouTube, with a Creative Common license.
16 papers · 1 benchmark
CPED (Chinese Personalized and Emotional Dialogue)
We construct a dataset named CPED from 40 Chinese TV shows.
15 papers · 3 benchmarks
CelebV-Text comprises 70,000 in-the-wild face video clips with diverse visual content, each paired with 20 texts generated using the proposed semi-automatic text generation strategy.
15 papers · 0 benchmarks
A large-scale video dataset, featuring clips from movies with detailed captions.
15 papers · 1 benchmark
Countix is a real world dataset of repetition videos collected in the wild (i.e.YouTube) covering a wide range of semantic settings with significant challenges such as camera and object motion, diverse set of periods and counts, and…
15 papers · 1 benchmark
Everybody Dance Now is a dataset of videos that can be used for training and motion transfer.
15 papers · 0 benchmarks
First-Person Hand Action Benchmark is a collection of RGB-D video sequences comprised of more than 100K frames of 45 daily hand action categories, involving 26 different objects in several hand configurations.
15 papers · 2 benchmarks
LIVE-FB LSVQ (LIVE-FB Large-Scale Social Video Quality (LSVQ) Database)
No-reference (NR) perceptual video quality assessment (VQA) is a complex, unsolved, and important problem to social and streaming media applications.
15 papers · 1 benchmark
MedVidQA (Medical Video Question Answering)
The MedVidQA dataset contains the collection of 3, 010 manually created health-related questions and timestamps as visual answers to those questions from trusted video sources, such as accredited medical schools with an established…
15 papers · 0 benchmarks
Large-scale American Sign Language (ASL) - English dataset collected from online video sites (e.g., YouTube).
15 papers · 0 benchmarks
PANDA is the first gigaPixel-level humAN-centric viDeo dAtaset, for large-scale, long-term, and multi-object visual analysis.
15 papers · 0 benchmarks
PointOdyssey is a large-scale synthetic dataset, and data generation framework, for the training and evaluation of long-term fine-grained tracking algorithms.
15 papers · 1 benchmark
A large-scale dataset for retrieval and event localisation in video.
15 papers · 1 benchmark
REDS (REalistic and Diverse Scenes dataset
realistic and dynamic scenes)
The realistic and dynamic scenes (REDS) dataset was proposed in the NTIRE19 Challenge.
15 papers · 1 benchmark
TITAN consists of 700 labeled video-clips (with odometry) captured from a moving vehicle on highly interactive urban traffic scenes in Tokyo.
15 papers · 0 benchmarks
4DFAB is a large scale database of dynamic high-resolution 3D faces which consists of recordings of 180 subjects captured in four different sessions spanning over a five-year period (2012 - 2017), resulting in a total of over 1,800,000 3D…
14 papers · 0 benchmarks
AVSD (Audio-Visual Scene-Aware Dialog)
The Audio Visual Scene-Aware Dialog (AVSD) dataset, or DSTC7 Track 3, is a audio-visual dataset for dialogue understanding.
14 papers · 1 benchmark
DeepStab is a dataset for online video stabilization consisting of synchronized steady/unsteady video pairs collected via a well designed hand-held hardware.
14 papers · 0 benchmarks
Fisheye dataset comprises of synthetically generated fisheye sequences and fisheye video sequences captured with an actual fisheye camera designed for fisheye motion estimation.
14 papers · 0 benchmarks
HAA500 (Human-Centric Atomic Action Dataset)
HAA500 is a manually annotated human-centric atomic action dataset for action recognition on 500 classes with over 591k labeled frames.
14 papers · 1 benchmark
The human-Related version of the ShanghaiTech Campus, was first presented by Morais et al.
14 papers · 1 benchmark
HiFiMask is a large-scale High-Fidelity Mask dataset, namely CASIA-SURF HiFiMask (briefly HiFiMask).
14 papers · 0 benchmarks
LAV-DF (Localized Audio Visual DeepFake Dataset)
Localized Audio Visual DeepFake Dataset (LAV-DF).
14 papers · 1 benchmark
LIVE-YT-HFR comprises of 480 videos having 6 different frame rates, obtained from 16 diverse contents.
14 papers · 1 benchmark
MGif is a dataset of videos containing movements of different cartoon animals.
14 papers · 1 benchmark
MMPD (Multi-Domain Mobile Video Physiology Dataset)
The Multi-domain Mobile Video Physiology Dataset (MMPD), comprising 11 hours(1152K frames) of recordings from mobile phones of 33 subjects.
14 papers · 0 benchmarks
OVBench is a benchmark tailored for real-time video understanding: - Memory, Perception, and Prediction of Temporal Contexts: Questions are framed to reference the present state of entities, requiring models to memorize/perceive/predict…
14 papers · 1 benchmark
RealEstate10K is a large dataset of camera poses corresponding to 10 million frames derived from about 80,000 video clips, gathered from about 10,000 YouTube videos.
14 papers · 1 benchmark
ViTT (Video Timeline Tags)
The ViTT dataset consists of human produced segment-level annotations for 8,169 videos.
14 papers · 2 benchmarks
The dataset contains 21 full-HD videos, each around 1 hr long, captured at six different locations.
13 papers · 1 benchmark
DNA-Rendering is a large-scale, high-fidelity repository of human performance data for neural actor rendering.
13 papers · 0 benchmarks
EVE (End-to-end Video-based Eye-tracking)
EVE (End-to-end Video-based Eye-tracking) is a dataset for eye-tracking.
13 papers · 0 benchmarks
The EgoDexter dataset provides both 2D and 3D pose annotations for 4 testing video sequences with 3190 frames.
13 papers · 0 benchmarks
Prophesee’s GEN1 Automotive Detection Dataset is the largest Event-Based Dataset to date.
13 papers · 1 benchmark
A video dataset for benchmarking upsampling methods.
13 papers · 0 benchmarks
In order to create the TED-talks dataset, 3,035 YouTube videos were downloaded using the "TED talks" query.
13 papers · 1 benchmark
This dataset encompasses a diverse range of tactile features that are instrumental in bifurcating various material properties.
13 papers · 0 benchmarks
Existing audio-visual event localization (AVE) handles manually trimmed videos with only a single instance in each of them.
13 papers · 1 benchmark
WSVD (Web Stereo Video Dataset)
The Web Stereo Video Dataset consists of 553 stereoscopic videos from YouTube.
13 papers · 0 benchmarks
The DramaQA focuses on two perspectives: 1) Hierarchical QAs as an evaluation metric based on the cognitive developmental stages of human intelligence.
12 papers · 1 benchmark
The Hands in action dataset (HIC) dataset has RGB-D sequences of hands interacting with objects.
12 papers · 0 benchmarks
HyperKvasir dataset contains 110,079 images and 374 videos where it captures anatomical landmarks and pathological and normal findings.
12 papers · 2 benchmarks
Neptune (Neptune Long Video Understanding Benchmark)
Neptune is a dataset consisting of challenging question-answer-decoy (QAD) sets for long videos (up to 15 minutes).
12 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.