Home › Datasets › modality › RGB Video
RGB Video datasets
archive 2025-07-28
87 datasets carry the modality tag "RGB Video", ordered by the archive's paper count. Page 2 of 2: 39 shown of 87. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
RGB Video datasets 49–87 of 87
PHANTOM (Physical Anomalous Trajectory or Motion (PHANTOM))
To evaluate the presented approaches, we created the Physical Anomalous Trajectory or Motion (PHANTOM) dataset consisting of six classes featuring everyday objects or physical setups, and showing nine different kinds of anomalies.
2 papers · 1 benchmark
PUMaVOS (Partial and Unusual Masks for Video Object Segmentation)
PUMaVOS is a dataset of challenging and practical use cases inspired by the movie production industry.
2 papers · 0 benchmarks
RHM (Rhm: Robot house multi-view human activity recognition dataset)
The Robot House Multi-View dataset (RHM) contains four views: Front, Back, Ceiling, and Robot Views.
2 papers · 1 benchmark
RL Unplugged is suite of benchmarks for offline reinforcement learning.
2 papers · 0 benchmarks
SB20 (Sugar Beet 2020 University of Bonn)
Video sequences captured at a field on Campus Kleinaltendorf (CKA), University of Bonn, captured by BonBot-I, an autonomous weeding robot.
2 papers · 0 benchmarks
SuperCaustics is a simulation tool made in Unreal Engine for generating massive computer vision datasets that include transparent objects.
2 papers · 0 benchmarks
TI1K Dataset (Thumb Index 1000 Hand & Fingertip Detection Dataset)
Thumb Index 1000 (TI1K) is a dataset of 1000 hand images with the hand bounding box, and thumb and index fingertip positions.
2 papers · 0 benchmarks
AzSLD (AzSLD - Azerbaijani Sign Language Dataset)
The Azerbaijani Sign Language Dataset (AzSLD) is a comprehensive, large dataset designed to facilitate the development and evaluation of machine learning models for the recognition and translation of Azerbaijani Sign Language (AzSL).
1 paper · 0 benchmarks
In this dataset two robots, Baxter and UR5, perform 8 behaviors (look, grasp, pick, hold, shake, lower, drop, and push) on 95 objects that vary by 5 color (blue, green, red, white, and yellow), 6 contents (wooden button, plastic dices,…
1 paper · 0 benchmarks
Bukva (Bukva: Russian Sign Language Alphabet)
We introduce a video dataset Bukva for Russian Dactyl Recognition task.
1 paper · 1 benchmark
The field of biomechanics is at a turning point, with marker-based motion capture set to be replaced by portable and inexpensive hardware, rapidly improving markerless tracking algorithms, and open datasets that will turn these new…
1 paper · 0 benchmarks
In this dataset an uppertorso humanoid robot with 7-DOF arm explored 100 different objects belonging to 20 different categories using 10 behaviors: Look, Crush, Grasp, Hold, Lift, Drop, Poke, Push, Shake and Tap.
1 paper · 0 benchmarks
ConSLAM (Construction Dataset for SLAM)
ConSLAM is a real-world dataset collected periodically on a construction site to measure the accuracy of mobile scanners' SLAM algorithms.
1 paper · 0 benchmarks
ConsInv is a stereo RGB + IMU dataset designed for Dynamic SLAM testing and contains two subsets: - ConsInv-Indoors contains sequences in an office setting where small objects are moved.
1 paper · 0 benchmarks
DADE (Driving Agents in Dynamic Environments)
The DADE dataset, short for Driving Agents in Dynamic Environments, is a synthetic dataset designed for the training and evaluation of methods for the task of semantic segmentation in the context of autonomous driving agents navigating…
1 paper · 0 benchmarks
FreeMan is the first large-scale multi-view human motion dataset under real scenarios.
1 paper · 0 benchmarks
HA-ViD (HA-ViD: A Human Assembly Video Dataset)
Understanding comprehensive assembly knowledge from videos is critical for futuristic ultra-intelligent industry.
1 paper · 0 benchmarks
InfraParis is a novel and versatile dataset supporting multiple tasks across three modalities: RGB, depth, and infrared.
1 paper · 0 benchmarks
Involves data where a robot interacts with 5.1 cm colored blocks to complete an order-fulfillment style block stacking task.
1 paper · 0 benchmarks
LSA-T (Lengua de Señas Argentina - Traducción)
LSA-T is the first continuous Argentinian Sign Language (LSA) dataset.
1 paper · 1 benchmark
MultiSenseBadminton (MultiSenseBadminton: Wearable Sensor–Based Biomechanical Dataset for Evaluation of Badminton Performance)
The sports industry is witnessing an increasing trend of utilizing multiple synchronized sensors for player data collection, enabling personalized training systems with multi-perspective real-time feedback.
1 paper · 0 benchmarks
Robot@Home2 (Robot@Home2, a robotic dataset of home environments)
Robot@Home2, is an enhanced version aimed at improving usability and functionality for developing and testing mobile robotics and computer vision algorithms.
1 paper · 0 benchmarks
The SoccerTrack dataset comprises top-view and wide-view video footage annotated with bounding boxes.
1 paper · 0 benchmarks
In this dataset UR5 robot used 6 tools: metal-scissor, metal-whisk, plastic-knife, plastic-spoon, wooden-chopstick, and wooden-fork to perform 6 behaviors: look, stirring-slow, stirring-fast, stirring-twist, whisk, and poke.
1 paper · 0 benchmarks
VETRA is a dataset for vehicle tracking in aerial image sequences and presents unique challenges such as low frame rates, small and fast-moving objects, as well as high camera movement.
1 paper · 0 benchmarks
The benchmark for VPData, the largest video inpainting dataset, which comprises over 390K clips (> 866.7 hours) and features precise masks and detailed video captions.
1 paper · 0 benchmarks
The largest video inpainting dataset comprises over 390K clips (> 866.7 hours), featuring precise masks and detailed video captions.
1 paper · 0 benchmarks
We provide separate training, development and test data.
1 paper · 0 benchmarks
WiFiCam dataset for through-wall imaging based on WiFi channel state information.
1 paper · 0 benchmarks
The YCB-Ev dataset contains synchronized RGB-D frames and event data that enables evaluating 6DoF object pose estimation algorithms using these modalities.
1 paper · 0 benchmarks
Dataset for the DREAMING - Diminished Reality for Emerging Applications in Medicine through Inpainting Challenge!
0 papers · 0 benchmarks
a large video dataset captured with UAVs in different complex real-world scenes, with multiple representations, suitable for multi-task learning.
0 papers · 0 benchmarks
HEADSET (HEADSET: Human Emotion Awareness under Partial Occlusions Multimodal DataSET)
The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications.
0 papers · 0 benchmarks
Overview The IITKGPFence dataset is designed for tasks related to fence-like occlusion detection, defocus blur, depth mapping, and object segmentation.
0 papers · 0 benchmarks
InfiniteRep is a synthetic, open-source dataset for fitness and physical therapy (PT) applications.
0 papers · 0 benchmarks
Infinity AI's Spills Basic Dataset is a synthetic, open-source dataset for safety applications.
0 papers · 0 benchmarks
L-SVD (Large-Scale Selfie Video Dataset (L-SVD): A Benchmark for Emotion Recognition)
Welcome to L-SVD L-SVD is an extensive and rigorously curated video dataset aimed at transforming the field of emotion recognition.
0 papers · 0 benchmarks
Human activity recognition and clinical biomechanics are challenging problems in physical telerehabilitation medicine.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.