Home › Datasets › modality › RGB Video

RGB Video datasets

archive 2025-07-28

87 datasets carry the modality tag "RGB Video", ordered by the archive's paper count. Page 1 of 2: 48 shown of 87. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

RGB Video datasets 1–48 of 87

Cambridge Landmarks, a large scale outdoor visual relocalisation dataset taken around Cambridge University.
109 papers · 0 benchmarks
CSL-Daily (Chinese Sign Language Corpus) is a large-scale continuous SLT dataset.
63 papers · 2 benchmarks
Avenue Dataset contains 16 training and 21 testing video clips.
48 papers · 2 benchmarks
UAV-Human is a large dataset for human behavior understanding with UAVs.
47 papers · 5 benchmarks
How2Sign (A Large-scale Multimodal Dataset for Continuous American Sign Language)
The How2Sign is a multimodal and multiview continuous American Sign Language (ASL) dataset consisting of a parallel corpus of more than 80 hours of sign language videos and a set of corresponding modalities including speech, English…
44 papers · 3 benchmarks
V2V4Real is a large-scale real-world multi-modal dataset for V2V perception.
35 papers · 0 benchmarks
EMDB contains in-the-wild videos of human activity recorded with a hand-held iPhone.
32 papers · 2 benchmarks
V2X-Sim, short for vehicle-to-everything simulation, is the a synthetic collaborative perception dataset in autonomous driving developed by AI4CE Lab at NYU and MediaBrain Group at SJTU to facilitate collaborative perception between…
32 papers · 1 benchmark
UVO (Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation)
UVO is a new benchmark for open-world class-agnostic object segmentation in videos.
27 papers · 2 benchmarks
VOID (Visual Odometry with Inertial and Depth)
The dataset was collected using the Intel RealSense D435i camera, which was configured to produce synchronized accelerometer and gyroscope measurements at 400 Hz, along with synchronized VGA-size (640 x 480) RGB and depth streams at 30 Hz.
26 papers · 1 benchmark
The Easy Communications (EasyCom) dataset is a world-first dataset designed to help mitigate the cocktail party effect from an augmented-reality (AR) -motivated multi-sensor egocentric world view.
22 papers · 4 benchmarks
STARSS23 (STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events)
The Sony-TAu Realistic Spatial Soundscapes 2023 (STARSS23) dataset contains multichannel recordings of sound scenes in various rooms and environments, together with temporal and spatial annotations of prominent events belonging to a set of…
21 papers · 0 benchmarks
BURST is a benchmark suite built upon TAO that requires tracking and segmenting multiple objects from camera video.
18 papers · 5 benchmarks
SCAND (Socially CompliAnt Navigation Dataset)
Have you wondered how autonomous mobile robots should share space with humans in public spaces?
18 papers · 0 benchmarks
Memorability dataset with 10000 3-second videos.
17 papers · 0 benchmarks
The SUN-SEG dataset is a high-quality per-frame annotated VPS dataset, which includes 158,690 frames from the famous SUN dataset.
17 papers · 1 benchmark
The SUN-SEG dataset is a high-quality per-frame annotated VPS dataset, which includes 158,690 frames from the famous SUN dataset.
17 papers · 1 benchmark
ViViD++ (Vision for Visibility Dataset)
A dataset capturing diverse visual data formats that target varying luminance conditions, and was recorded from alternative vision sensors, by handheld or mounted on a car, repeatedly in the same space but in different conditions.
17 papers · 0 benchmarks
A database with 2,000 videos captured by surveillance cameras in real-world scenes.
16 papers · 1 benchmark
CCVID (Clothes-Changing Video person re-ID)
Clothes-Changing Video person re-ID (CCVID) is a dataset constructed from the raw data of a gait recognition dataset, i.e.
15 papers · 1 benchmark
RECON (RECON Outdoor Navigation Dataset)
https://sites.google.com/view/recon-robot/dataset
12 papers · 0 benchmarks
4D-OR includes a total of 6734 scenes, recorded by six calibrated RGB-D Kinect sensors 1 mounted to the ceiling of the OR, with one frame-per-second, providing synchronized RGB and depth images.
11 papers · 3 benchmarks
HANDAL (HANDAL: A Dataset of Real-World Manipulable Object Categories with Pose Annotations, Affordances, and Reconstructions)
We present the HANDAL dataset for category-level object pose estimation and affordance prediction.
8 papers · 0 benchmarks
CHAD (Charlotte Anomaly Dataset)
CHAD: Charlotte Anomaly Dataset CHAD is high-resolution, multi-camera dataset for surveillance video anomaly detection.
7 papers · 1 benchmark
UBI-Fights (Abnormal Event Detection Dataset)
UBI-Fights - Concerning a specific anomaly detection and still providing a wide diversity in fighting scenarios, the UBI-Fights dataset is a unique new large-scale dataset of 80 hours of video fully annotated at the frame level.
7 papers · 2 benchmarks
WebLINX (Real-World Website Navigation with Multi-Turn)
WebLINX is a large-scale benchmark of 100K interactions across 2300 expert demonstrations of conversational web navigation.
6 papers · 1 benchmark
A new dataset with significant occlusions related to object manipulation.
6 papers · 0 benchmarks
BUP20 (Sweet Pepper 2020 University of Bonn)
Video sequences from a glasshouse environment in Campus Kleinaltendorf(CKA), University of Bonn, captured by PATHoBot, a glasshouse monitoring robot.
5 papers · 0 benchmarks
DurLAR (A High-Fidelity 128-Channel LiDAR Dataset with Panoramic Ambient and Reflectivity Imagery)
DurLAR is a high-fidelity 128-channel 3D LiDAR dataset with panoramic ambient (near infrared) and reflectivity imagery for multi-modal autonomous driving applications.
5 papers · 0 benchmarks
A large-scale multi-modal dataset to facilitate research and studies that concentrate on vision-wireless systems.
5 papers · 1 benchmark
HuPR (Human Pose with Millimeter Wave Radar)
HuPR is a human pose estimation benchmark is created using cross-calibrated mmWave radar sensors and a monocular RGB camera for cross-modality training of radar-based human pose estimation.
4 papers · 0 benchmarks
MMToM-QA (Multimodal Theory of Mind Question Answering)
MMToM-QA is the first multimodal benchmark to evaluate machine Theory of Mind (ToM), the ability to understand people's minds.
4 papers · 0 benchmarks
This is a dataset for video deinterlacing problem.
4 papers · 1 benchmark
Something-Something-100 is a dataset split created from Something-Something V2.
4 papers · 1 benchmark
SportsPose (SportsPose - A Dynamic 3D sports pose dataset)
Accurate 3D human pose estimation is essential for sports analytics, coaching, and injury prevention.
4 papers · 0 benchmarks
The dataset is designed specifically to solve a range of computer vision problems (2D-3D tracking, posture) faced by biologists while designing behavior studies with animals.
3 papers · 0 benchmarks
SoccerNet-GSR (SoccerNet Game State Reconstruction)
The SoccerNet Game State Reconstruction task is a novel high level computer vision task that is specific to sports analytics.
3 papers · 0 benchmarks
TTStroke-21 ME21 (TTStroke-21 for MediaEval 2021)
This task offers researchers an opportunity to test their fine-grained classification methods for detecting and recognizing strokes in table tennis videos.
3 papers · 2 benchmarks
UESTC-MMEA-CL (A multi-modal egocentric activity dataset for continual learning)
UESTC-MMEA-CL is a new multi-modal activity dataset for continual egocentric activity recognition, which is proposed to promote future studies on continual learning for first-person activity recognition in wearable applications.
3 papers · 0 benchmarks
VBR (VBR: A Vision Benchmark in Rome)
This dataset presents a vision and perception research dataset collected in Rome, featuring RGB data, 3D point clouds, IMU, and GPS data.
3 papers · 0 benchmarks
3DYoga90 (3DYoga90: A Hierarchical Video Dataset for Yoga Pose Understanding)
3DYoga90 is organized within a three-level label hierarchy.
2 papers · 0 benchmarks
A real-world dataset, with hyper-accurate digital counterpart & comprehensive ground-truth annotation.
2 papers · 1 benchmark
CholecT40 (Cholecystectomy Action Triplet)
CholecT40 is the first endoscopic dataset introduced to enable research on fine-grained action recognition in laparoscopic surgery.
2 papers · 1 benchmark
Dynamic OLAT Dataset (ShanghaiTech MARS Dynamic OLAT Dataset)
To provide ground truth supervision for video consistency modeling, we build up a high-quality dynamic OLAT dataset.
2 papers · 0 benchmarks
Fetoscopic Placental Vessel Segmentation and Registration (FetReg2021) challenge was organized as part of the MICCAI2021 Endoscopic Vision (EndoVis) challenge.
2 papers · 0 benchmarks
LoTE-Animal (LoTE-Animal: A Long Time-span Dataset for Endangered Animal Behavior Understanding)
Understanding and analyzing animal behavior is increasingly essential to protect endangered animal species.
2 papers · 1 benchmark
LuViRA (Lund University Vision, Radio, and Audio)
The Lund University Vision, Radio, and Audio (LuViRA) positioning dataset consists of 89 trajectories that are recorded in the Lund University Humanities Lab's Motion Capture (Mocap) Studio using a MIR200 robot as the targeted platform.
2 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.