Browse › Computer Vision › Monocular Depth Estimation › MIX-6
MIX-6 Benchmark (Monocular Depth Estimation)
Monocular Depth Estimation is the task of estimating the depth value (distance relative to the camera) of each pixel given a single (monocular) RGB image. This challenging task is a key prerequisite for determining scene understanding for applications such as 3D scene reconstruction, autonomous driving, and AR. State-of-the-art methods usually fall into one of two categories: designing a complex network that is powerful enough to directly regress the depth map, or splitting the input into bins or windows to reduce computational complexity. The most popular benchmarks are the KITTI and NYUv2 datasets. Models are typically evaluated using RMSE or absolute relative error.
The archive carries no text for this table; the description above is the archive's text for the task Monocular Depth Estimation. archive 2025-07-28
Results archive 2025-07-28
No rows in the archive for this table at snapshot 2025-07-28. It declares 1 metric (Zero-shot transfer) but no result was ever recorded against it. That says nothing about whether results exist elsewhere.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections