Home › Datasets › modality › RGB-D
RGB-D datasets
archive 2025-07-28
190 datasets carry the modality tag "RGB-D", ordered by the archive's paper count. Page 1 of 4: 48 shown of 190. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets
RGB-D datasets 1–48 of 190
ScanNet is an instance-level indoor RGB-D dataset that includes both 2D and 3D data.
1,595 papers · 21 benchmarks
The NYU-Depth V2 data set is comprised of video sequences from a variety of indoor scenes as recorded by both the RGB and Depth cameras from the Microsoft Kinect.
986 papers · 16 benchmarks
The SUN RGBD dataset contains 10335 real RGB-D images of room scenes.
477 papers · 11 benchmarks
NTU RGB+D is a large-scale dataset for RGB-D human action recognition.
476 papers · 9 benchmarks
The Matterport3D dataset is a large RGB-D dataset for scene understanding in indoor environments.
461 papers · 4 benchmarks
TUM RGB-D is an RGB-D dataset.
235 papers · 1 benchmark
SUNCG is a large-scale dataset of synthetic 3D scenes with dense volumetric annotations.
186 papers · 0 benchmarks
ALFRED (Action Learning From Realistic Environments and Directives)
ALFRED (Action Learning From Realistic Environments and Directives), is a new benchmark for learning a mapping from natural language instructions and egocentric vision to sequences of actions for household tasks.
173 papers · 0 benchmarks
The YCB-Video dataset is a large-scale video dataset for 6D object pose estimation.
164 papers · 5 benchmarks
SUN3D contains a large-scale RGB-D video database, with 8 annotated sequences.
126 papers · 0 benchmarks
ETHD is a multi-view stereo benchmark / 3D reconstruction benchmark that covers a variety of indoor and outdoor scenes.
121 papers · 3 benchmarks
For many fundamental scene understanding tasks, it is difficult or impossible to obtain per-pixel ground truth labels from real images.
108 papers · 4 benchmarks
T-LESS is a dataset for estimating the 6D pose, i.e.
94 papers · 2 benchmarks
DIODE (Dense Indoor and Outdoor Depth)
Diode Dense Indoor/Outdoor DEpth (DIODE) is the first standard dataset for monocular depth estimation comprising diverse indoor and outdoor scenes acquired with the same hardware setup.
86 papers · 2 benchmarks
ARKitScenes is an RGB-D dataset captured with the widely available Apple LiDAR scanner.
75 papers · 2 benchmarks
REAL275 is a benchmark for category-level pose estimation.
67 papers · 1 benchmark
The CAD-60 and CAD-120 data sets comprise of RGB-D video sequences of humans performing activities which are recording using the Microsoft Kinect sensor.
65 papers · 1 benchmark
SceneNN is an RGB-D scene dataset consisting of more than 100 indoor scenes.
63 papers · 1 benchmark
The EYEDIAP dataset is a dataset for gaze estimation from remote RGB, and RGB-D (standard vision and depth), cameras.
57 papers · 3 benchmarks
BEHAVE is a full body human-object interaction dataset with multi-view RGBD frames and corresponding 3D SMPL and object fits along with the annotated contacts between them.
53 papers · 3 benchmarks
UAV-Human is a large dataset for human behavior understanding with UAVs.
47 papers · 5 benchmarks
The ScanNet200 benchmark studies 200-class 3D semantic segmentation - an order of magnitude more class categories than previous 3D scene understanding benchmarks.
45 papers · 3 benchmarks
How2Sign (A Large-scale Multimodal Dataset for Continuous American Sign Language)
The How2Sign is a multimodal and multiview continuous American Sign Language (ASL) dataset consisting of a parallel corpus of more than 80 hours of sign language videos and a set of corresponding modalities including speech, English…
44 papers · 3 benchmarks
Dataset containing RGB-D data of 4 large scenes, comprising a total of 12 rooms, for the purpose of RGB and RGB-D camera relocalization.
41 papers · 0 benchmarks
A hand-object interaction dataset with 3D pose annotations of hand and object.
36 papers · 2 benchmarks
We build a large-scale, comprehensive, and high-quality synthetic dataset for city-scale neural rendering researches.
33 papers · 0 benchmarks
WMCA (Wide Multi Channel Presentation Attack)
The Wide Multi Channel Presentation Attack (WMCA) database consists of 1941 short video recordings of both bonafide and presentation attacks from 72 different identities.
33 papers · 1 benchmark
InteriorNet is a RGB-D for large scale interior scene understanding and mapping.
30 papers · 0 benchmarks
AVD (Active Vision Dataset)
AVD focuses on simulating robotic vision tasks in everyday indoor environments using real imagery.
29 papers · 1 benchmark
GraspNet-1Billion provides large-scale training data and a standard evaluation platform for the task of general robotic grasping.
29 papers · 1 benchmark
OCID (Object Clutter Indoor Dataset)
Developing robot perception systems for handling objects in the real-world requires computer vision algorithms to be carefully scrutinized with respect to the expected operating domain.
29 papers · 1 benchmark
SBU-Kinect-Interaction dataset version 2.0 comprises of RGB-D video sequences of humans performing interaction activities that are recording using the Microsoft Kinect sensor.
28 papers · 4 benchmarks
THuman2.0 Dataset contains 500 high-quality human scans captured by a dense DLSR rig.
28 papers · 1 benchmark
The Drive&Act dataset is a state of the art multi modal benchmark for driver behavior recognition.
26 papers · 1 benchmark
VOID (Visual Odometry with Inertial and Depth)
The dataset was collected using the Intel RealSense D435i camera, which was configured to produce synchronized accelerometer and gyroscope measurements at 400 Hz, along with synchronized VGA-size (640 x 480) RGB and depth streams at 30 Hz.
26 papers · 1 benchmark
REALY (Region-aware benchmark based on the LYHM)
The REALY benchmark aims to introduce a region-aware evaluation pipeline to measure the fine-grained normalized mean square error (NMSE) of 3D face reconstruction methods from under-controlled image sets.
24 papers · 2 benchmarks
ReDWeb (Relative Depth from Web)
The ReDWeb dataset consists of 3600 RGB-RD image pairs collected from the Web.
22 papers · 0 benchmarks
SCAND (Socially CompliAnt Navigation Dataset)
Have you wondered how autonomous mobile robots should share space with humans in public spaces?
18 papers · 0 benchmarks
CDTB (Color-and-Depth Tracking)
Source: https://www.vicos.si/Projects/CDTB 4.2 State-of-the-art Comparison A TH CTB (color-and-depth visual object tracking) dataset is recorded by several passive and active RGB-D setups and contains indoor as well as outdoor sequences…
17 papers · 0 benchmarks
ViViD++ (Vision for Visibility Dataset)
A dataset capturing diverse visual data formats that target varying luminance conditions, and was recorded from alternative vision sensors, by handheld or mounted on a car, repeatedly in the same space but in different conditions.
17 papers · 0 benchmarks
First-Person Hand Action Benchmark is a collection of RGB-D video sequences comprised of more than 100K frames of 45 daily hand action categories, involving 26 different objects in several hand configurations.
15 papers · 2 benchmarks
Washington RGB-D is a widely used testbed in the robotic community, consisting of 41,877 RGB-D images organized into 300 instances divided in 51 classes of common indoor objects (e.g.
15 papers · 0 benchmarks
This dataset accompanies our paper on synthesizing the 3D Ken Burns effect from a single image.
13 papers · 0 benchmarks
The EgoDexter dataset provides both 2D and 3D pose annotations for 4 testing video sequences with 3190 frames.
13 papers · 0 benchmarks
CIRCLE is a dataset containing 10 hours of full-body reaching motion from 5 subjects across nine scenes, paired with ego-centric information of the environment represented in various forms, such as RGBD videos.
12 papers · 1 benchmark
The Hands in action dataset (HIC) dataset has RGB-D sequences of hands interacting with objects.
12 papers · 0 benchmarks
HRWSI (High-Resolution Web Stereo Image)
The HRWSI dataset consists of about 21K diverse high-resolution RGB-D image pairs derived from the Web stereo images.
12 papers · 0 benchmarks
4D-OR includes a total of 6734 scenes, recorded by six calibrated RGB-D Kinect sensors 1 mounted to the ceiling of the OR, with one frame-per-second, providing synchronized RGB and depth images.
11 papers · 3 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.