Home › Datasets › modality › RGB-D
RGB-D datasets
archive 2025-07-28
190 datasets carry the modality tag "RGB-D", ordered by the archive's paper count. Page 2 of 4: 48 shown of 190. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
RGB-D datasets 49–96 of 190
GSL (Greek Sign Language)
Dataset Description The Greek Sign Language (GSL) is a large-scale RGB+D dataset, suitable for Sign Language Recognition (SLR) and Sign Language Translation (SLT).
11 papers · 1 benchmark
RGB-D-D is a large-scale dataset for depth map super-resolution (SR).
11 papers · 0 benchmarks
Dynamic Replica is a synthetic dataset of stereo videos featuring humans and animals in virtual environments.
10 papers · 0 benchmarks
Bonn RGB-D Dynamic is a dataset for RGB-D SLAM, containing highly dynamic sequences.
9 papers · 1 benchmark
HQ-WMCA (High-Quality Wide Multi-Channel Attack database)
The High-Quality Wide Multi-Channel Attack database (HQ-WMCA) database consists of 2904 short multi-modal video recordings of both bona-fide and presentation attacks.
9 papers · 0 benchmarks
DeepLoc is a large-scale urban outdoor localization dataset.
8 papers · 0 benchmarks
HANDAL (HANDAL: A Dataset of Real-World Manipulable Object Categories with Pose Annotations, Affordances, and Reconstructions)
We present the HANDAL dataset for category-level object pose estimation and affordance prediction.
8 papers · 0 benchmarks
X-Humans consists of 20 subjects (11 males, 9 females) with various clothing types and hair style.
8 papers · 0 benchmarks
The 7-Scenes dataset is a collection of tracked RGB-D camera frames.
7 papers · 0 benchmarks
ChangeSim is a dataset aimed at online scene change detection (SCD) and more.
7 papers · 2 benchmarks
The HandNet dataset contains depth images of 10 participants' hands non-rigidly deforming in front of a RealSense RGB-D camera.
7 papers · 0 benchmarks
The MMBody dataset provides human body data with motion capture, GT mesh, Kinect RGBD, and millimeter wave sensor data.
7 papers · 0 benchmarks
This is a 3D action recognition dataset, also known as 3D Action Pairs dataset.
7 papers · 1 benchmark
NTU RGB+D 2D is a curated version of NTU RGB+D often used for skeleton-based action prediction and synthesis.
7 papers · 1 benchmark
The SynthHands dataset is a dataset for hand pose estimation which consists of real captured hand motion retargeted to a virtual hand with natural backgrounds and interactions with different objects.
7 papers · 0 benchmarks
TransCG is the first large-scale real-world dataset for transparent object depth completion and grasping, which contains 57,715 RGB-D images of 51 transparent objects and many opaque objects captured from different perspectives (~240…
7 papers · 1 benchmark
This dataset contains 70 (30 falls + 40 activities of daily living) sequences.
7 papers · 0 benchmarks
The Bimanual Actions Dataset is a collection of 540 RGB-D videos, showing subjects perform bimanual actions in a kitchen or workshop context.
6 papers · 0 benchmarks
The Freiburg Forest dataset was collected using a Viona autonomous mobile robot platform equipped with cameras for capturing multi-spectral and multi-modal images.
6 papers · 2 benchmarks
IndustReal (IndustReal Dataset of Egocentric Videos for Procedure Understanding)
IndustReal is an ego-centric, multi-modal dataset where 27 participants are challenged to perform assembly and maintenance procedures on a construction-toy car.
6 papers · 3 benchmarks
Generate high-quality 3D ground-truth shapes for reconstruction evaluation is extremely challenging because even 3D scanners can only generate pseudo ground-truth shapes with artefacts.
6 papers · 0 benchmarks
A new dataset with significant occlusions related to object manipulation.
6 papers · 0 benchmarks
A Large Dataset of Object Scans is a dataset of more than ten thousand 3D scans of real objects.
5 papers · 0 benchmarks
BUP20 (Sweet Pepper 2020 University of Bonn)
Video sequences from a glasshouse environment in Campus Kleinaltendorf(CKA), University of Bonn, captured by PATHoBot, a glasshouse monitoring robot.
5 papers · 0 benchmarks
COME15K is an RGB-D saliency detection dataset which contains 15,625 image pairs with high quality polygon-/scribble-/object-/instance-/rank-level annotations.
5 papers · 0 benchmarks
DIMO (Dataset of Industrial Metal Objects)
The Industrial Metal Objects dataset is a diverse dataset of industrial metal objects.
5 papers · 0 benchmarks
HARPER (Exploring 3D Human Pose Estimation and Forecasting from the Robot’s Perspective: The HARPER Dataset)
We introduce HARPER, a novel dataset for 3D body pose estimation and forecast in dyadic interactions between users and \spot, the quadruped robot manufactured by Boston Dynamics.
5 papers · 3 benchmarks
NERDS 360 (NeRF for Reconstruction, Decomposition and Scene Synthesis of 360° outdoor scenes)
We present a large-scale dataset for 3D urban scene understanding.
5 papers · 0 benchmarks
SERV-CT (SERV-CT: A disparity dataset from CT for validation of endoscopic 3D reconstruction)
Endoscopic stereo reconstruction for surgical scenes gives rise to specific problems, including the lack of clear corner features, highly specular surface properties, and the presence of blood and smoke.
5 papers · 0 benchmarks
A large-scale multi-modal dataset to facilitate research and studies that concentrate on vision-wireless systems.
5 papers · 1 benchmark
Provides a large-scale synthetic dataset which contains accurate ground truth depth of various photo-realistic scenes.
4 papers · 0 benchmarks
DoMSEV (Dataset of Multimodal Semantic Egocentric Video)
The Dataset of Multimodal Semantic Egocentric Video (DoMSEV) contains 80-hours of multimodal (RGB-D, IMU, and GPS) data related to First-Person Videos with annotations for recorder profile, frame scene, activities, interaction, and…
4 papers · 0 benchmarks
EDEN (Enclosed garDEN) is a multimodal synthetic dataset, a dataset for nature-oriented applications.
4 papers · 0 benchmarks
FINO-Net is a multimodal (RGB, depth and audio) dataset, containing 229 real-world manipulation data of 5 different manipulation types recorded with a Baxter robot.
4 papers · 0 benchmarks
FewSOL (A Dataset for Few-Shot Object Learning in Robotic Environments)
The Few-Shot Object Learning (FewSOL) dataset can be used for object recognition with a few images per object.
4 papers · 0 benchmarks
MMToM-QA (Multimodal Theory of Mind Question Answering)
MMToM-QA is the first multimodal benchmark to evaluate machine Theory of Mind (ToM), the ability to understand people's minds.
4 papers · 0 benchmarks
MUAD (Multiple Uncertainties for Autonomous Driving)
The MUAD dataset (Multiple Uncertainties for Autonomous Driving), consisting of 10,413 realistic synthetic images with diverse adverse weather conditions (night, fog, rain, snow), out-of-distribution objects, and annotations for semantic…
4 papers · 0 benchmarks
MUSES offers 2500 multi-modal scenes, evenly distributed across various combinations of weather conditions (clear, fog, rain, and snow) and types of illumination (daytime, nighttime).
4 papers · 5 benchmarks
From PARIS: Part-level Reconstruction and Motion Analysis for Articulated Objects: 5.1.
4 papers · 0 benchmarks
TRansPose is a large-scale multispectral dataset that combines stereo RGB-D, TIR (TIR) images, and object poses to promote transparent object research.
4 papers · 0 benchmarks
AnoVox is a large-scale benchmark for ANOmaly detection in autonomous driving.
3 papers · 0 benchmarks
The Composable activities dataset consists of 693 videos that contain activities in 16 classes performed by 14 actors.
3 papers · 0 benchmarks
Operating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient safety.
3 papers · 3 benchmarks
UASOL (A large-scale high-resolution outdoor stereo dataset)
The UASOL an RGB-D stereo dataset, that contains 160902 frames, filmed at 33 different scenes, each with between 2 k and 10 k frames.
3 papers · 1 benchmark
The UAVA,UAV-Assistant, dataset is specifically designed for fostering applications which consider UAVs and humans as cooperative agents.
3 papers · 0 benchmarks
A synthetic depth estimation dataset for benchmark rendered from a high-quality CAD indoor environment - About 3.5K RGBD pairs with left-right stereo - Challenging viewing direction - Challenging different light condition
3 papers · 1 benchmark
This work was undertaken by members of the Lincoln Centre for Autonomous Systems, University of Lincoln, UK.
2 papers · 0 benchmarks
Depth vision has been recently used in many locomotion devices with the objective to ease the life of disabled people toward reaching more ecological lifestyle.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.