Home › Datasets › modality › Videos
Videos datasets
archive 2025-07-28
1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 13 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
Videos datasets 577–624 of 1,014
Content4All is a collection of six open research datasets aimed at automatic sign language translation research.
3 papers · 0 benchmarks
Countix-AV is a dataset for repetitive action counting by sight and sound created by repurposing the Countix dataset.
3 papers · 0 benchmarks
DCASE2014 is an audio classification benchmark.
3 papers · 0 benchmarks
DFDM (Deepfake videos generated from different models)
We created a new dataset, named DFDM, with 6,450 Deepfake videos generated by different Autoencoder models.
3 papers · 0 benchmarks
DIVOTrack is a cross-view multi-object tracking dataset for DIVerse Open scenes with dense tracking pedestrians in realistic and non-experimental environments.
3 papers · 0 benchmarks
The Deep Fakes Dataset is a collection of "in the wild" portrait videos for deepfake detection.
3 papers · 0 benchmarks
DeepFake MNIST+ is a deepfake facial animation dataset.
3 papers · 0 benchmarks
From Grounded Human-Object Interaction Hotspots from Video (ICCV'19): We collect annotations for interaction keypoints on EPIC Kitchens in order to quantitatively evaluate our method in parallel to the OPRA dataset (where annotations are…
3 papers · 1 benchmark
The BIWI Walking Pedestrians dataset consists of walking pedestrians in busy scenarios from a birds eye view.
3 papers · 0 benchmarks
Ego4D-HCap is a hierarchical video captioning dataset comprised of a three-tier hierarchy of captions: short clip-level captions, medium-length video segment descriptions, and long-range video-level summaries.
3 papers · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
3 papers · 0 benchmarks
Fitness-AQA (Fitness Action Quality Assessment [ECCV 2022])
Largest, first-of-its-kind, in-the-wild, fine-grained workout/exercise posture analysis dataset, covering three different exercises: BackSquat, Barbell Row, and Overhead Press.
3 papers · 0 benchmarks
Ford Campus Vision and Lidar Data Set is a dataset collected by an autonomous ground vehicle testbed, based upon a modified Ford F-250 pickup truck.
3 papers · 0 benchmarks
FunQA is a challenging video question answering (QA) dataset specifically designed to evaluate and enhance the depth of video reasoning based on counter-intuitive and fun videos.
3 papers · 0 benchmarks
GraSP (Holistic and Multi-Granular Surgical Scene Understanding of Prostatectomies)
Holistic and Multi-Granular Surgical Scene Understanding of Prostatectomies (GraSP) dataset, a curated benchmark that models surgical scene understanding as a hierarchy of complementary tasks with varying levels of granularity.
3 papers · 1 benchmark
GroOT (Grounded Multiple Object Tracking)
One of the recent trends in vision problems is to use natural language captions to describe the objects of interest.
3 papers · 0 benchmarks
HDM05 is a MoCap (motion capture) dataset.
3 papers · 1 benchmark
HOPE-Video (Household Objects for Pose Estimation)
The HOPE-Video dataset contains 10 video sequences (2038 frames) with 5-20 objects on a tabletop scene captured by a robot arm-mounted RealSense D415 RGBD camera.
3 papers · 0 benchmarks
IfAct (Identifying Human Actions Visible in Online Vlogs)
We consider the task of identifying human actions visible in online videos.
3 papers · 0 benchmarks
This Dataset consists of 2120 sequences of binary masks of pedestrians.
3 papers · 1 benchmark
The dataset contains the annotations of characters' visual appearances, in the form of tracks of face bounding boxes, and the associations with characters' textual mentions, when available.
3 papers · 1 benchmark
MERL Shopping is a dataset for training and testing action detection algorithms.
3 papers · 0 benchmarks
MISP2021 (Multimodal Information Based Speech Processing 2021)
The MISP2021 challenge dataset is a collection of audio-visual conversational data recorded in a home TV scenario using distant multi-microphones.
3 papers · 0 benchmarks
Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety.
3 papers · 3 benchmarks
MMDB (Multimodal Dyadic Behavior)
Multimodal Dyadic Behavior (MMDB) dataset is a unique collection of multimodal (video, audio, and physiological) recordings of the social and communicative behavior of toddlers.
3 papers · 0 benchmarks
MOMA-LRG (Multi-Object Multi-Actor activity parsing with Language-Refined Graphs)
A dataset dedicated to multi-object, multi-actor activity parsing.
3 papers · 1 benchmark
The pioneering eyeblink detection dataset is characterized by three key features: (1) Sample with multi-human instances.
3 papers · 2 benchmarks
MSR-VTT Adverbs is a subset from MSR-VTT with extracted verb-adverb annotations.
3 papers · 2 benchmarks
Frame-to-frame video alignment/synchronization
3 papers · 1 benchmark
MeLa BitChute is a near-complete dataset of over 3M videos from 61K channels over 2.5 years (June 2019 to December 2021) from the social video hosting platform BitChute, a commonly used alternative to YouTube.
3 papers · 0 benchmarks
Replay data from human players and AI agents navigating in a 3D game environment.
3 papers · 0 benchmarks
OREBA (Objectively Recognizing Eating Behavior and Associated Intake)
The OREBA dataset aims to provide a comprehensive multi-sensor recording of communal intake occasions for researchers interested in automatic detection of intake gestures.
3 papers · 0 benchmarks
OpenLane-V2 is the world's first perception and reasoning benchmark for scene structure in autonomous driving.
3 papers · 1 benchmark
OSAI introduces OpenTTGames - an open dataset aimed at evaluation of different computer vision tasks in Table Tennis: ball detection, semantic segmentation of humans, table and scoreboard and fast in-game events spotting.
3 papers · 0 benchmarks
PETRAW (PEg TRAnsfer Workflow recognition by different modalities)
PETRAW data set was composed of 150 sequences of peg transfer training sessions.
3 papers · 6 benchmarks
A large-scale video portrait dataset that contains 291 videos from 23 conference scenes with 14K frames.
3 papers · 0 benchmarks
Polyps in the colon are widely known cancer precursors identified by colonoscopy.
3 papers · 1 benchmark
QST contains 1,167 video clips that are cut out from 216 time-lapse 4K videos collected from YouTube, which can be used for a variety of tasks, such as (high-resolution) video generation, (high-resolution) video prediction,…
3 papers · 0 benchmarks
RLV (Reinforcement Learning with Videos)
We provide video observations of humans performing two simple tasks in natural environments.
3 papers · 0 benchmarks
RoboBEV is a robustness evaluation benchmark tailored for camera-based bird's eye view (BEV) perception under natural data corruptions and domain shift.
3 papers · 0 benchmarks
SDN (Situated Dialogue Navigation)
Situated Dialogue Navigation (SDN) is a navigation benchmark of 183 trials with a total of 8415 utterances, around 18.7 hours of control streams, and 2.9 hours of trimmed audio.
3 papers · 0 benchmarks
A short clip of video may contain progression of multiple events and an interesting story line.
3 papers · 3 benchmarks
The SoccerNet Game State Reconstruction task is a novel high level computer vision task that is specific to sports analytics.
3 papers · 0 benchmarks
SurgT is a dataset for benchmarking 2D Trackers in Minimally Invasive Surgery (MIS).
3 papers · 0 benchmarks
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips,…
3 papers · 1 benchmark
The Tongue and Lips (TaL) corpus is a multi-speaker corpus of ultrasound images of the tongue and video images of lips.
3 papers · 0 benchmarks
The TbV dataset is large-scale dataset created to allow the community to improve the state of the art in machine learning tasks related to mapping, that are vital for self-driving.
3 papers · 0 benchmarks
Tragic Talkers is an audio-visual dataset consisting of excerpts from the "Romeo and Juliet" drama captured with microphone arrays and multiple co-located cameras for light-field video.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.