Home › Datasets › task › Monocular Depth Estimation

Monocular Depth Estimation datasets

archive 2025-07-28

33 datasets carry the task tag "Monocular Depth Estimation" (the task itself: Monocular Depth Estimation), ordered by the archive's paper count. Page 1 of 1: 33 shown of 33. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Monocular Depth Estimation datasets 1–33 of 33

Cityscapes is a large-scale database which focuses on semantic understanding of urban street scenes.
3,702 papers · 51 benchmarks
KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute) is one of the most popular datasets for use in mobile robotics and autonomous driving.
3,661 papers · 137 benchmarks
NYUv2 (NYU-Depth V2)
The NYU-Depth V2 data set is comprised of video sequences from a variety of indoor scenes as recorded by both the RGB and Depth cameras from the Microsoft Kinect.
986 papers · 16 benchmarks
The Matterport3D dataset is a large RGB-D dataset for scene understanding in indoor environments.
461 papers · 4 benchmarks
Virtual KITTI is a photo-realistic synthetic video dataset designed to learn and evaluate computer vision models for several video understanding tasks: object detection and multi-object tracking, scene-level and instance-level semantic…
133 papers · 0 benchmarks
The Make3D dataset is a monocular Depth Estimation dataset that contains 400 single training RGB and depth map pairs, and 134 test samples.
129 papers · 1 benchmark
ETHD is a multi-view stereo benchmark / 3D reconstruction benchmark that covers a variety of indoor and outdoor scenes.
121 papers · 3 benchmarks
For many fundamental scene understanding tasks, it is difficult or impossible to obtain per-pixel ground truth labels from real images.
108 papers · 4 benchmarks
DIODE (Dense Indoor and Outdoor Depth)
Diode Dense Indoor/Outdoor DEpth (DIODE) is the first standard dataset for monocular depth estimation comprising diverse indoor and outdoor scenes acquired with the same hardware setup.
86 papers · 2 benchmarks
DDAD (Dense Depth for Autonomous Driving)
DDAD is a new autonomous driving benchmark from TRI (Toyota Research Institute) for long range (up to 250m) and dense depth estimation in challenging and diverse urban conditions.
73 papers · 1 benchmark
The Middlebury 2014 dataset contains a set of 23 high resolution stereo pairs for which known camera calibration parameters and ground truth disparity maps obtained with a structured light scanner are available.
59 papers · 2 benchmarks
Virtual KITTI 2 is an updated version of the well-known Virtual KITTI dataset which consists of 5 sequence clones from the KITTI tracking benchmark.
53 papers · 2 benchmarks
IBims-1 (Independent benchmark images and matched scans v1)
iBims-1 (independent Benchmark images and matched scans - version 1) is a new high-quality RGB-D dataset, especially designed for testing single-image depth estimation (SIDE) methods.
34 papers · 2 benchmarks
The MannequinChallenge Dataset (MQC) provides in-the-wild videos of people in static poses while a hand-held camera pans around the scene.
28 papers · 0 benchmarks
ReDWeb (Relative Depth from Web)
The ReDWeb dataset consists of 3600 RGB-RD image pairs collected from the Web.
22 papers · 0 benchmarks
Collects high quality 360 datasets with ground truth depth annotations, by re-using recently released large scale 3D datasets and re-purposing them to 360 via rendering.
18 papers · 0 benchmarks
This dataset accompanies our paper on synthesizing the 3D Ken Burns effect from a single image.
13 papers · 0 benchmarks
Detecting vehicles and representing their position and orientation in the three dimensional space is a key technology for autonomous driving.
13 papers · 3 benchmarks
WSVD (Web Stereo Video Dataset)
The Web Stereo Video Dataset consists of 553 stereoscopic videos from YouTube.
13 papers · 0 benchmarks
HRWSI (High-Resolution Web Stereo Image)
The HRWSI dataset consists of about 21K diverse high-resolution RGB-D image pairs derived from the Web stereo images.
12 papers · 0 benchmarks
An in-the-wild stereo image dataset, comprising 49,368 image pairs contributed by users of the Holopix mobile social platform.
12 papers · 0 benchmarks
HUMAN4D is a large and multimodal 4D dataset that contains a variety of human activities simultaneously captured by a professional marker-based MoCap, a volumetric capture and an audio recording system.
9 papers · 0 benchmarks
Are current 3D object tracking methods truely robust enough for low-fidelity depth sensors like the iPhone LiDAR?
8 papers · 2 benchmarks
Mid-Air, The Montefiore Institute Dataset of Aerial Images and Records, is a multi-purpose synthetic dataset for low altitude drone flights.
6 papers · 2 benchmarks
SYNS-Patches dataset, which is a subset of SYNS.
5 papers · 0 benchmarks
MUAD (Multiple Uncertainties for Autonomous Driving)
The MUAD dataset (Multiple Uncertainties for Autonomous Driving), consisting of 10,413 realistic synthetic images with diverse adverse weather conditions (night, fog, rain, snow), out-of-distribution objects, and annotations for semantic…
4 papers · 0 benchmarks
UASOL (A large-scale high-resolution outdoor stereo dataset)
The UASOL an RGB-D stereo dataset, that contains 160902 frames, filmed at 33 different scenes, each with between 2 k and 10 k frames.
3 papers · 1 benchmark
A synthetic depth estimation dataset for benchmark rendered from a high-quality CAD indoor environment - About 3.5K RGBD pairs with left-right stereo - Challenging viewing direction - Challenging different light condition
3 papers · 1 benchmark
VBR (VBR: A Vision Benchmark in Rome)
This dataset presents a vision and perception research dataset collected in Rome, featuring RGB data, 3D point clouds, IMU, and GPS data.
3 papers · 0 benchmarks
InfraParis is a novel and versatile dataset supporting multiple tasks across three modalities: RGB, depth, and infrared.
1 paper · 0 benchmarks
The dataset includes polarimetric, RGB and depth automotive (on the road) data.
1 paper · 0 benchmarks
SCARED-C (SCARED-Corrupted)
The dataset SCARED-C is introduced in the context of assessing robustness in endoscopic depth prediction models.
1 paper · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.