Home › Datasets › modality › Videos

Videos datasets

archive 2025-07-28

1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 9 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Videos datasets 385–432 of 1,014

Amazon Mechanical Turk (AMT) is used to collect annotations on HowTo100M videos.
9 papers · 0 benchmarks
We describe the 2020 edition of the DeepMind Kinetics human action dataset, which replenishes and extends the Kinetics-700 dataset.
9 papers · 1 benchmark
KnowIT VQA is a video dataset with 24,282 human-generated question-answer pairs about The Big Bang Theory.
9 papers · 0 benchmarks
MSASL is a real-life large-scale sign language data set comprising over 25,000 annotated videos.
9 papers · 1 benchmark
This is a dataset for a video inverse-tone-mapping task.
9 papers · 1 benchmark
MVK (Marine Video Kit)
The dataset contains single-shot videos taken from moving cameras in underwater environments.
9 papers · 1 benchmark
MuSe-CaR (Multimodal Sentiment Analysis in Car Reviews)
The MuSe-CAR database is a large, multimodal (video, audio, and text) dataset which has been gathered in-the-wild with the intention of further understanding Multimodal Sentiment Analysis in-the-wild, e.g., the emotional engagement that…
9 papers · 0 benchmarks
PSI-AVA is a dataset designed for holistic surgical scene understanding.
9 papers · 0 benchmarks
PTB-TIR is a Thermal InfraRed (TIR) pedestrian tracking benchmark, which provides 60 TIR sequences with mannuly annoations.
9 papers · 0 benchmarks
PathTrack is a dataset for person tracking which contains more than 15,000 person trajectories in 720 sequences.
9 papers · 0 benchmarks
Most existing MOT datasets are captured using pinhole cameras, which are characterized by a narrow-FoV and linear sensor motion.
9 papers · 1 benchmark
TEMPO (Localizing Moments in Video with Temporal Language)
TEMPOral reasoning in video and language (TEMPO) is a dataset that consists of two parts: a dataset with real videos and template sentences (TEMPO - Template Language) which allows for controlled studies on temporal language, and a human…
9 papers · 0 benchmarks
TUM-GAID (TUM Gait from Audio, Image and Depth) collects 305 subjects performing two walking trajectories in an indoor environment.
9 papers · 0 benchmarks
Thai-Chi-HD is a high resolution dataset which can be used as reference benchmark for evaluating frameworks for image animation and video generation.
9 papers · 3 benchmarks
The UCF Sports dataset consists of a set of actions collected from various sports which are typically featured on broadcast television channels such as the BBC and ESPN.
9 papers · 1 benchmark
UDIVA is a new non-acted dataset of face-to-face dyadic interactions, where interlocutors perform competitive and collaborative tasks with different behavior elicitation and cognitive workload.
9 papers · 0 benchmarks
VIL-100 is a video instance lane detection dataset, which contains 100 videos with in total 10,000 frames, acquired from different real traffic scenarios.
9 papers · 0 benchmarks
VLM2-Bench (VLM²-Bench)
VLM²-Bench: Benchmarking Vision-Language Models on Visual Cue Matching Description VLM²-Bench is the first comprehensive benchmark designed to evaluate vision-language models' (VLMs) ability to visually link matching cues across…
9 papers · 1 benchmark
VOT2020 is a Visual Object Tracking benchmark for short-term tracking in RGB.
9 papers · 1 benchmark
The Zenseact Open Dataset (ZOD) is a large-scale and diverse multi-modal autonomous driving (AD) dataset, created by researchers at Zenseact.
9 papers · 0 benchmarks
BosphorusSign22k is a benchmark dataset for vision-based user-independent isolated Sign Language Recognition (SLR).
8 papers · 0 benchmarks
Are current 3D object tracking methods truely robust enough for low-fidelity depth sensors like the iPhone LiDAR?
8 papers · 2 benchmarks
DarkTrack2021 is a challenging nighttime UAV tracking benchmark, which contains 110 challenging sequences with over 100 K frames in total.
8 papers · 0 benchmarks
FVI (Free-form Video Inpainting)
The Free-Form Video Inpainting dataset is a dataset used for training and evaluation video inpainting models.
8 papers · 0 benchmarks
HANDAL (HANDAL: A Dataset of Real-World Manipulable Object Categories with Pose Annotations, Affordances, and Reconstructions)
We present the HANDAL dataset for category-level object pose estimation and affordance prediction.
8 papers · 0 benchmarks
Kinetics-100 is a dataset split created from the Kinetics dataset to evaluate the performance of few-shot action recognition models.
8 papers · 1 benchmark
MISAW (MIcro-Surgical Anastomose Workflow recognition on training sessions)
The MISAW data set is composed of 27 sequences of micro-surgical anastomosis on artificial blood vessels performed by 3 surgeons and 3 engineering students.
8 papers · 1 benchmark
OPERAnet is a multimodal activity recognition dataset acquired from radio frequency and vision-based sensors.
8 papers · 0 benchmarks
OPRA (Online Product Reviews for Affordances)
The OPRA Dataset was introduced in Demo2Vec: Reasoning Object Affordances From Online Videos (CVPR'18) for reasoning object affordances from online demonstration videos.
8 papers · 2 benchmarks
The Oulu-NPU face presentation attack detection database consists of 4950 real access and attack videos.
8 papers · 1 benchmark
OpenS2V-Eval introduces 180 prompts from seven major categories of S2V, which incorporate both real and synthetic test data.
8 papers · 1 benchmark
ROBUST-MIS (Robust Medical Instrument Segmentation Challenge 2019)
The ROBUST-MIS dataset was made available to support the Robust Medical Instrument Segmentation (ROBUST-MIS) Challenge 2019, part of the Endoscopic Vision Challenge associated with MICCAI.
8 papers · 1 benchmark
The SAMM Long Videos dataset consists of 147 long videos with 343 macro-expressions and 159 micro-expressions.
8 papers · 0 benchmarks
SEND (Stanford Emotional Narratives Dataset)
SEND (Stanford Emotional Narratives Dataset) is a set of rich, multimodal videos of self-paced, unscripted emotional narratives, annotated for emotional valence over time.
8 papers · 0 benchmarks
TikTok Dataset (Learning High Fidelity Depths of Dressed Humans by Watching Social Media Dance Videos)
We learn high fidelity human depths by leveraging a collection of social media dance videos scraped from the TikTok mobile social networking application.
8 papers · 0 benchmarks
The Video-based Multimodal Summarization with Multimodal Output (VMSMO) corpus consists of 184,920 document-summary pairs, with 180,000 training pairs, 2,460 validation and test pairs.
8 papers · 0 benchmarks
A new dataset describing textual stories for events.
8 papers · 0 benchmarks
VideoXum is an enriched large-scale dataset for cross-modal video summarization.
8 papers · 1 benchmark
The YouTube-100M data set consists of 100 million YouTube videos: 70M training videos, 10M evaluation videos, and 20M validation videos.
8 papers · 0 benchmarks
YouTube-ASL is a large-scale, open-domain corpus of American Sign Language (ASL) videos and accompanying English captions drawn from YouTube.
8 papers · 0 benchmarks
ACAV100M (Automatically Curated Audio-Visual)
ACAV100M processes 140 million full-length videos (total duration 1,030 years) which are used to produce a dataset of 100 million 10-second clips (31 years) with high audio-visual correspondence.
7 papers · 0 benchmarks
The Atari Grand Challenge dataset is a large dataset of human Atari 2600 replays.
7 papers · 0 benchmarks
BS-RSC is a real-world rolling shutter (RS) correction dataset and a corresponding model to correct the RS frames in a distorted video.
7 papers · 1 benchmark
CHAD (Charlotte Anomaly Dataset)
CHAD: Charlotte Anomaly Dataset CHAD is high-resolution, multi-camera dataset for surveillance video anomaly detection.
7 papers · 1 benchmark
ChangeSim is a dataset aimed at online scene change detection (SCD) and more.
7 papers · 2 benchmarks
The Human Related version of UBnormal ("UBnormal: New Benchmark for Supervised Open-Set Video Anomaly Detection," Acsintoae et al.) was introduced by Flaborea et al.
7 papers · 1 benchmark
HiREST (HIerarchical REtrieval and STep-captioning)
HiREST (HIerarchical REtrieval and STep-captioning) dataset is a benchmark that covers hierarchical information retrieval and visual/textual stepwise summarization from an instructional video corpus.
7 papers · 0 benchmarks
The KUMC dataset for polyp detection and classification was collected from the University of Kansas Medical Center.
7 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.