Home › Datasets › modality › Videos

Videos datasets

archive 2025-07-28

1,014 datasets carry the modality tag "Videos", ordered by the archive's paper count. Page 19 of 22: 48 shown of 1,014. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Videos datasets 865–912 of 1,014

In this Pre-Contest Workshop Video Recordings folder: Seven screen and audio recordings of seven pre-contest workshops
1 paper · 0 benchmarks
QoEVAVE (Quality of Experience Evaluation of Interactive Virtual Environments with Audiovisual Scenes)
Quality of Experience Evaluation of Interactive Virtual Environments with Audiovisual Scenes (QoEVAVE) provides an initial audiovisual database consiting of 12 sequences capturing real-life nature and urban scenes.
1 paper · 0 benchmarks
In this dataset, various objects are arranged on a white table.
1 paper · 0 benchmarks
RTB (Robot Tracking Benchmark)
The Robot Tracking Benchmark (RTB) is a synthetic dataset that facilitates the quantitative evaluation of 3D tracking algorithms for multi-body objects.
1 paper · 1 benchmark
Reactive Diffusion Policy-Dataset (Dataset of Reactive Diffusion Policy)
Two versions of the dataset are offered: one is the full dataset used to train the models in our paper, and the other is a mini dataset for easier examination.
1 paper · 0 benchmarks
A real-world stereo video dataset, containing 1200 frame pairs with real-world color and sharpness mismatches caused by beam splitter.
1 paper · 0 benchmarks
The Reasonable Crowd dataset is a dataset to evaluate autonomous driving in a limited operating domain.
1 paper · 0 benchmarks
Replay is a collection of multi-view, multi-modal videos of humans interacting socially.
1 paper · 0 benchmarks
Robot@Home2 (Robot@Home2, a robotic dataset of home environments)
Robot@Home2, is an enhanced version aimed at improving usability and functionality for developing and testing mobile robotics and computer vision algorithms.
1 paper · 0 benchmarks
SARA motion (Synthetic Actors and Real Actions)
Sara motion is a 3D motion dataset, named Synthetic Actors and Real Actions (SARA), for training a model to produce motion embeddings suitable for reasoning about motion similarity.
1 paper · 0 benchmarks
A Chinese sign language dataset that includes dialogue information.
1 paper · 0 benchmarks
SF20K (Short-Films 20K)
Short-Films 20K (SF20K) is the largest publicly available movie dataset.
1 paper · 0 benchmarks
The dataset SFU-HW-Objects-v1 contains bounding boxes and object class labels for High Efficiency Video Coding (HEVC) v1 Common Test Conditions (CTC) video sequences.
1 paper · 0 benchmarks
SFU-HW-Tracks is a dataset for Object Tracking on raw video sequences that contains object annotations with unique object identities (IDs) for the High Efficiency Video Coding (HEVC) v1 Common Test Conditions (CTC) sequences.
1 paper · 0 benchmarks
A dataset collected in a set of experiments that involves human participants and a robot.
1 paper · 0 benchmarks
SOMPT22 (Surveillance Oriented Multi-Pedestrian Tracking Dataset (SOMPT22))
SOMPT22 is a multi-object tracking (MOT) benchmark focused on surveillance-style pedestrian tracking.
1 paper · 0 benchmarks
SOTVerse is a user-defined task space of single object tracking.
1 paper · 0 benchmarks
APPROVE consists of curated YouTube videos annotated with educational content.
1 paper · 1 benchmark
SSv2-Spatio-Temporal (Something Someting v2-Spatio-Temporal)
We use Something-Something v2 dataset to obtain the generation prompts and ground truth masks from real action videos.
1 paper · 0 benchmarks
A large-scale Japanese video caption dataset consisting of 79,822 videos and 399,233 captions.
1 paper · 0 benchmarks
STVD-PVCD (Partial Video Copy Detection Dataset)
STVD is the largest public dataset on the PVCD task.
1 paper · 1 benchmark
SaGA (The Bielefeld Speech and Gesture Alignment Corpus (SaGA))
The primary data of the SaGA corpus are made up of 25 dialogs of interlocutors (50), who engage in a spatial communication task combining direction-giving and sight description.
1 paper · 0 benchmarks
Please refer to the Zenodo page for a detailed description: https://zenodo.org/records/15665101
1 paper · 0 benchmarks
Procedural videos show step-by-step demonstrations of tasks like recipe preparation.
1 paper · 0 benchmarks
The dataset consists of 90 000 grayscale videos that show two objects of equal shape and size in which one object approaches the other one.
1 paper · 0 benchmarks
Sintel 4D LFV (Sintel 4D Light Field Video Dataset)
A medium-scale synthetic 4D Light Field video dataset for depth (disparity) estimation.
1 paper · 0 benchmarks
We introduce a large-scale video dataset Slovo for Russian Sign Language task.
1 paper · 1 benchmark
SoccerNet-Echoes (SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset)
SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset.
1 paper · 0 benchmarks
The SoccerTrack dataset comprises top-view and wide-view video footage annotated with bounding boxes.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Songdo Traffic (Songdo Traffic: High Accuracy Georeferenced Vehicle Trajectories from a Large-Scale Study in a Smart City)
The Songdo Traffic dataset delivers precisely georeferenced vehicle trajectories captured through high-altitude bird's-eye view (BeV) drone footage over Songdo International Business District, South Korea.
1 paper · 0 benchmarks
We collect a dataset of 805 clean videos that show the action of pouring water in a container.
1 paper · 1 benchmark
Surgical Hands is a dataset that provides multi-instance articulated hand pose annotations for in-vivo videos.
1 paper · 0 benchmarks
The dataset is collected from the Youtube videos that contains fight instances in it.
1 paper · 0 benchmarks
SynoClip Dataset The SynoClip dataset is a comprehensive and standard dataset specifically designed for the video synopsis task.
1 paper · 0 benchmarks
Synthetic dataset comprising three different environments for multi-camera dynamic novel view synthesis for soccer.
1 paper · 0 benchmarks
Our dataset augments the TAO dataset with amodal bounding box annotations for fully invisible, out-of-frame, and occluded objects.
1 paper · 0 benchmarks
The TED VCR Video Retrieval Dataset is a multimodal collection derived from publicly available TED Talks.
1 paper · 0 benchmarks
THGP (Temporal Hands Guns and Phones Dataset)
Temporal Hands Guns and Phones (THGP) dataset, is a collection of 5960 video frames (5000 for training and 960 for testing).
1 paper · 0 benchmarks
TLFM dataset (TLFM dataset for microscopy image sequence generation)
TLFM dataset structured in sequences of at least nine timesteps.
1 paper · 0 benchmarks
The dataset has been designed to represent true web videos in the wild, with good visual quality and diverse content characteristics, The test video collection for TRECVID-AVS2019-TRECVID-AVS2021, which contains 1,082,649 web video clips,…
1 paper · 1 benchmark
The dataset is composed of 100 video sequences densely annotated with 60K bounding boxes, 17 sequence attributes, 13 action verb attributes and 29 target object attributes.
1 paper · 0 benchmarks
TUMTraffic-VideoQA is a novel dataset designed to understand spatiotemporal video in complex roadside traffic scenarios.
1 paper · 0 benchmarks
TVPReid (Text-to-Video Person Re-identification)
The TVPReid dataset contains 6559 pedestrian videos, each of which is annotated with two text descriptions, for a total of 13118 descriptions.
1 paper · 0 benchmarks
TYC Dataset (The TYC Dataset for Understanding Instance-Level Semantics and Motions of Cells in Microstructures)
We introduce the trapped yeast cell (TYC) dataset, a novel dataset for understanding instance-level semantics and motions of cells in microstructures.
1 paper · 0 benchmarks
The RBO dataset of articulated objects and interactions is a collection of 358 RGB-D video sequences (67:18 minutes) of humans manipulating 14 articulated objects under varying conditions (light, perspective, background, interaction).
1 paper · 0 benchmarks
The TimberVision dataset consists of more than 2k annotated RGB images and contains a total of 51k trunk components including cut and lateral surfaces, thereby surpassing any existing dataset in this domain in terms of both quantity and…
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.