Home › Datasets › task › Depth Estimation
Depth Estimation datasets
archive 2025-07-28
77 datasets carry the task tag "Depth Estimation" (the task itself: Depth Estimation), ordered by the archive's paper count. Page 1 of 2: 48 shown of 77. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Depth Estimation datasets 1–48 of 77
Cityscapes is a large-scale database which focuses on semantic understanding of urban street scenes.
3,702 papers · 51 benchmarks
KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute) is one of the most popular datasets for use in mobile robotics and autonomous driving.
3,661 papers · 137 benchmarks
ScanNet is an instance-level indoor RGB-D dataset that includes both 2D and 3D data.
1,595 papers · 21 benchmarks
The NYU-Depth V2 data set is comprised of video sequences from a variety of indoor scenes as recorded by both the RGB and Depth cameras from the Microsoft Kinect.
986 papers · 16 benchmarks
The Matterport3D dataset is a large RGB-D dataset for scene understanding in indoor environments.
461 papers · 4 benchmarks
TUM RGB-D is an RGB-D dataset.
235 papers · 1 benchmark
The Middlebury Stereo dataset consists of high-resolution stereo sequences with complex geometry and pixel-accurate ground-truth disparity data.
223 papers · 5 benchmarks
SUNCG is a large-scale dataset of synthetic 3D scenes with dense volumetric annotations.
186 papers · 0 benchmarks
The MegaDepth dataset is a dataset for single-view depth prediction that includes 196 different locations reconstructed from COLMAP SfM/MVS.
152 papers · 0 benchmarks
The 2D-3D-S dataset provides a variety of mutually registered modalities from 2D, 2.5D and 3D domains, with instance-level semantic and geometric annotations.
147 papers · 6 benchmarks
Taskonomy provides a large and high-quality dataset of varied indoor scenes.
147 papers · 2 benchmarks
Virtual KITTI is a photo-realistic synthetic video dataset designed to learn and evaluate computer vision models for several video understanding tasks: object detection and multi-object tracking, scene-level and instance-level semantic…
133 papers · 0 benchmarks
The Make3D dataset is a monocular Depth Estimation dataset that contains 400 single training RGB and depth map pairs, and 134 test samples.
129 papers · 1 benchmark
SUN3D contains a large-scale RGB-D video database, with 8 annotated sequences.
126 papers · 0 benchmarks
ETHD is a multi-view stereo benchmark / 3D reconstruction benchmark that covers a variety of indoor and outdoor scenes.
121 papers · 3 benchmarks
For many fundamental scene understanding tasks, it is difficult or impossible to obtain per-pixel ground truth labels from real images.
108 papers · 4 benchmarks
DIODE (Dense Indoor and Outdoor Depth)
Diode Dense Indoor/Outdoor DEpth (DIODE) is the first standard dataset for monocular depth estimation comprising diverse indoor and outdoor scenes acquired with the same hardware setup.
86 papers · 2 benchmarks
DDAD (Dense Depth for Autonomous Driving)
DDAD is a new autonomous driving benchmark from TRI (Toyota Research Institute) for long range (up to 250m) and dense depth estimation in challenging and diverse urban conditions.
73 papers · 1 benchmark
The Middlebury 2014 dataset contains a set of 23 high resolution stereo pairs for which known camera calibration parameters and ground truth disparity maps obtained with a structured light scanner are available.
59 papers · 2 benchmarks
Virtual KITTI 2 is an updated version of the well-known Virtual KITTI dataset which consists of 5 sequence clones from the KITTI tracking benchmark.
53 papers · 2 benchmarks
DrivingStereo contains over 180k images covering a diverse set of driving scenarios, which is hundreds of times larger than the KITTI Stereo dataset.
50 papers · 0 benchmarks
DENSE (Depth Estimation oN Synthetic Events)
DENSE (Depth Estimation oN Synthetic Events) is a new dataset with synthetic events and perfect ground truth.
49 papers · 1 benchmark
The MannequinChallenge Dataset (MQC) provides in-the-wild videos of people in static poses while a hand-held camera pans around the scene.
28 papers · 0 benchmarks
OASIS (Open Annotations of Single Image Surfaces)
A dataset for single-image 3D in the wild consisting of annotations of detailed 3D geometry for 140,000 images.
28 papers · 3 benchmarks
2D-3D Match Dataset is a new dataset of 2D-3D correspondences by leveraging the availability of several 3D datasets from RGB-D scans.
27 papers · 0 benchmarks
VOID (Visual Odometry with Inertial and Depth)
The dataset was collected using the Intel RealSense D435i camera, which was configured to produce synchronized accelerometer and gyroscope measurements at 400 Hz, along with synchronized VGA-size (640 x 480) RGB and depth streams at 30 Hz.
26 papers · 1 benchmark
ReDWeb (Relative Depth from Web)
The ReDWeb dataset consists of 3600 RGB-RD image pairs collected from the Web.
22 papers · 0 benchmarks
Collects high quality 360 datasets with ground truth depth annotations, by re-using recently released large scale 3D datasets and re-purposing them to 360 via rendering.
18 papers · 0 benchmarks
The KITTI-Depth dataset includes depth maps from projected LiDAR point clouds that were matched against the depth estimation from the stereo cameras.
14 papers · 0 benchmarks
WSVD (Web Stereo Video Dataset)
The Web Stereo Video Dataset consists of 553 stereoscopic videos from YouTube.
13 papers · 0 benchmarks
HRWSI (High-Resolution Web Stereo Image)
The HRWSI dataset consists of about 21K diverse high-resolution RGB-D image pairs derived from the Web stereo images.
12 papers · 0 benchmarks
An in-the-wild stereo image dataset, comprising 49,368 image pairs contributed by users of the Holopix mobile social platform.
12 papers · 0 benchmarks
Depth in the Wild is a dataset for single-image depth perception in the wild, i.e., recovering depth from a single image taken in unconstrained settings.
11 papers · 0 benchmarks
Dynamic Replica is a synthetic dataset of stereo videos featuring humans and animals in virtual environments.
10 papers · 0 benchmarks
HUMAN4D is a large and multimodal 4D dataset that contains a variety of human activities simultaneously captured by a professional marker-based MoCap, a volumetric capture and an audio recording system.
9 papers · 0 benchmarks
Are current 3D object tracking methods truely robust enough for low-fidelity depth sensors like the iPhone LiDAR?
8 papers · 2 benchmarks
TikTok Dataset (Learning High Fidelity Depths of Dressed Humans by Watching Social Media Dance Videos)
We learn high fidelity human depths by leveraging a collection of social media dance videos scraped from the TikTok mobile social networking application.
8 papers · 0 benchmarks
The Stanford Light Field Archive is a collection of several light fields for research in computer graphics and vision.
7 papers · 0 benchmarks
DurLAR (A High-Fidelity 128-Channel LiDAR Dataset with Panoramic Ambient and Reflectivity Imagery)
DurLAR is a high-fidelity 128-channel 3D LiDAR dataset with panoramic ambient (near infrared) and reflectivity imagery for multi-modal autonomous driving applications.
5 papers · 0 benchmarks
The Middlebury 2006 is a stereo dataset of indoor scenes with multiple handcrafted layouts.
5 papers · 0 benchmarks
NERDS 360 (NeRF for Reconstruction, Decomposition and Scene Synthesis of 360° outdoor scenes)
We present a large-scale dataset for 3D urban scene understanding.
5 papers · 0 benchmarks
SERV-CT (SERV-CT: A disparity dataset from CT for validation of endoscopic 3D reconstruction)
Endoscopic stereo reconstruction for surgical scenes gives rise to specific problems, including the lack of clear corner features, highly specular surface properties, and the presence of blood and smoke.
5 papers · 0 benchmarks
SYNS-Patches dataset, which is a subset of SYNS.
5 papers · 0 benchmarks
Provides a large-scale synthetic dataset which contains accurate ground truth depth of various photo-realistic scenes.
4 papers · 0 benchmarks
CocoDoom is a collection of pre-recorded data extracted from Doom gaming sessions along with annotations in the MS Coco format.
4 papers · 0 benchmarks
The DCM dataset is composed of 772 annotated images from 27 golden age comic books.
4 papers · 3 benchmarks
EDEN (Enclosed garDEN) is a multimodal synthetic dataset, a dataset for nature-oriented applications.
4 papers · 0 benchmarks
The endoscopic SLAM dataset (EndoSLAM) is a dataset for depth estimation approach for endoscopic videos.
4 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.