{"url":"/task/monocular-depth-estimation","name":"Monocular Depth Estimation","slug":"monocular-depth-estimation","description_markdown":"**Monocular Depth Estimation** is the task of estimating the depth value (distance relative to the camera) of each pixel given a single (monocular) RGB image. This challenging task is a key prerequisite for determining scene understanding for applications such as 3D scene reconstruction, autonomous driving, and AR. State-of-the-art methods usually fall into one of two categories: designing a complex network that is powerful enough to directly regress the depth map, or splitting the input into bins or windows to reduce computational complexity.  The most popular benchmarks are the KITTI and NYUv2 datasets. Models are typically evaluated using RMSE or absolute relative error. \r\n\r\n<span class=\"description-source\">Source: [Defocus Deblurring Using Dual-Pixel Data ](https://arxiv.org/abs/2005.00305)</span>","categories":[{"name":"Computer Vision","url":"/area/computer-vision"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":876,"papers_with_code":430,"benchmarks":24,"benchmark_tables_in_archive":25,"benchmark_tables_shown":25,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":33,"subtasks":0,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/monocular-depth-estimation-on-nyu-depth-v2","slug":"monocular-depth-estimation-on-nyu-depth-v2","dataset":"NYU-Depth V2","dataset_url":"/dataset/nyuv2","rows_in_archive":85,"metrics":["absolute relative error","RMSE","log 10","Delta < 1.25","Delta < 1.25^2","Delta < 1.25^3"],"first_row_in_archive_order":{"model":"HybridDepth","paper_title":"HybridDepth: Robust Metric Depth Fusion by Leveraging Depth from Focus and Single-Image Priors","paper_url":"/paper/hybriddepth-robust-depth-fusion-for-mobile-ar","paper_date":"2024-07-26","arxiv_id":"2407.18443","code_links":[{"title":"cake-lab/hybriddepth","url":"https://github.com/cake-lab/hybriddepth"}],"syntology":{"n":7,"n_ran":6,"n_unverified":1,"n_pointer_only":7}}},{"leaderboard":"/sota/monocular-depth-estimation-on-kitti-eigen","slug":"monocular-depth-estimation-on-kitti-eigen","dataset":"KITTI Eigen split","dataset_url":"/dataset/kitti","rows_in_archive":79,"metrics":["absolute relative error","RMSE","Sq Rel","RMSE log","Delta < 1.25","Delta < 1.25^2","Delta < 1.25^3","Square relative error (SqRel)"],"first_row_in_archive_order":{"model":"SPIDepth","paper_title":"SPIdepth: Strengthened Pose Information for Self-supervised Monocular Depth Estimation","paper_url":"/paper/spidepth-strengthened-pose-information-for","paper_date":"2024-04-18","arxiv_id":"2404.12501","code_links":[{"title":"Lavreniuk/SPIdepth","url":"https://github.com/Lavreniuk/SPIdepth"}],"syntology":{"n":14,"n_ran":13,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/monocular-depth-estimation-on-kitti-eigen-1","slug":"monocular-depth-estimation-on-kitti-eigen-1","dataset":"KITTI Eigen split unsupervised","dataset_url":"/dataset/kitti","rows_in_archive":55,"metrics":["absolute relative error","RMSE","RMSE log","Sq Rel","Delta < 1.25","Delta < 1.25^2","Delta < 1.25^3","Resolution","Mono","Test frames"],"first_row_in_archive_order":{"model":"SPIdepth","paper_title":"SPIdepth: Strengthened Pose Information for Self-supervised Monocular Depth Estimation","paper_url":"/paper/spidepth-strengthened-pose-information-for","paper_date":"2024-04-18","arxiv_id":"2404.12501","code_links":[{"title":"Lavreniuk/SPIdepth","url":"https://github.com/Lavreniuk/SPIdepth"}],"syntology":{"n":14,"n_ran":13,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/monocular-depth-estimation-on-eth3d","slug":"monocular-depth-estimation-on-eth3d","dataset":"ETH3D","dataset_url":"/dataset/eth3d","rows_in_archive":10,"metrics":["Delta < 1.25","absolute relative error"],"first_row_in_archive_order":{"model":"Distill Any Depth","paper_title":"Distill Any Depth: Distillation Creates a Stronger Monocular Depth Estimator","paper_url":"/paper/distill-any-depth-distillation-creates-a","paper_date":"2025-02-26","arxiv_id":"2502.19204","code_links":[{"title":"Westlake-AGI-Lab/Distill-Any-Depth","url":"https://github.com/Westlake-AGI-Lab/Distill-Any-Depth"}],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-nyu-depth-v2-4","slug":"monocular-depth-estimation-on-nyu-depth-v2-4","dataset":"NYU-Depth V2 self-supervised","dataset_url":"/dataset/nyuv2","rows_in_archive":8,"metrics":["Root mean square error (RMSE)","Absolute relative error (AbsRel)","delta_1","delta_2","delta_3"],"first_row_in_archive_order":{"model":"IndoorDepth","paper_title":"Deeper into Self-Supervised Monocular Indoor Depth Estimation","paper_url":"/paper/deeper-into-self-supervised-monocular-indoor","paper_date":"2023-12-03","arxiv_id":"2312.01283","code_links":[{"title":"fcntes/indoordepth","url":"https://github.com/fcntes/indoordepth"}],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-make3d","slug":"monocular-depth-estimation-on-make3d","dataset":"Make3D","dataset_url":"/dataset/make3d","rows_in_archive":6,"metrics":["Abs Rel","RMSE","Sq Rel"],"first_row_in_archive_order":{"model":"SPIDepth","paper_title":"SPIdepth: Strengthened Pose Information for Self-supervised Monocular Depth Estimation","paper_url":"/paper/spidepth-strengthened-pose-information-for","paper_date":"2024-04-18","arxiv_id":"2404.12501","code_links":[{"title":"Lavreniuk/SPIdepth","url":"https://github.com/Lavreniuk/SPIdepth"}],"syntology":{"n":14,"n_ran":13,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/monocular-depth-estimation-on-mid-air-dataset","slug":"monocular-depth-estimation-on-mid-air-dataset","dataset":"Mid-Air Dataset","dataset_url":"/dataset/mid-air-dataset","rows_in_archive":6,"metrics":["Abs Rel","RMSE","RMSE log","SQ Rel"],"first_row_in_archive_order":{"model":"M4Depth+U","paper_title":"A technique to jointly estimate depth and depth uncertainty for unmanned aerial vehicles","paper_url":"/paper/a-technique-to-jointly-estimate-depth-and","paper_date":"2023-05-31","arxiv_id":"2305.19780","code_links":[{"title":"michael-fonder/m4depthu","url":"https://github.com/michael-fonder/m4depthu"}],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-ddad","slug":"monocular-depth-estimation-on-ddad","dataset":"DDAD","dataset_url":"/dataset/ddad","rows_in_archive":4,"metrics":["RMSE","RMSE log","Sq Rel","absolute relative error","Delta < 1.25"],"first_row_in_archive_order":{"model":"AFNet","paper_title":"Adaptive Fusion of Single-View and Multi-View Depth for Autonomous Driving","paper_url":"/paper/adaptive-fusion-of-single-view-and-multi-view","paper_date":"2024-03-12","arxiv_id":"2403.07535","code_links":[{"title":"junda24/afnet","url":"https://github.com/junda24/afnet"}],"syntology":{"n":8,"n_ran":7,"n_unverified":1,"n_pointer_only":0}}},{"leaderboard":"/sota/monocular-depth-estimation-on-ibims-1","slug":"monocular-depth-estimation-on-ibims-1","dataset":"IBims-1","dataset_url":"/dataset/ibims-1","rows_in_archive":4,"metrics":["ORD","D3R","RMSE","δ1.25","absolute relative error"],"first_row_in_archive_order":{"model":"Miangoleh et al. (SGR)","paper_title":"Boosting Monocular Depth Estimation Models to High-Resolution via Content-Adaptive Multi-Resolution Merging","paper_url":"/paper/boosting-monocular-depth-estimation-models-to","paper_date":"2021-05-28","arxiv_id":"2105.14021","code_links":[{"title":"compphoto/BoostingMonocularDepth","url":"https://github.com/compphoto/BoostingMonocularDepth"}],"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}}},{"leaderboard":"/sota/monocular-depth-estimation-on-scared-c","slug":"monocular-depth-estimation-on-scared-c","dataset":"SCARED-C","dataset_url":"/dataset/scared-c","rows_in_archive":4,"metrics":["mDERS"],"first_row_in_archive_order":{"model":"AF-SfMLearner","paper_title":"EndoDepth: A Benchmark for Assessing Robustness in Endoscopic Depth Prediction","paper_url":"/paper/endodepth-a-benchmark-for-assessing","paper_date":"2024-09-30","arxiv_id":"2409.19930","code_links":[{"title":"Ivanrs297/endoscopycorruptions","url":"https://github.com/Ivanrs297/endoscopycorruptions"}],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-cityscapes","slug":"monocular-depth-estimation-on-cityscapes","dataset":"Cityscapes","dataset_url":"/dataset/cityscapes","rows_in_archive":3,"metrics":["RMSE","RMSE log","Absolute relative error (AbsRel)","Square relative error (SqRel)"],"first_row_in_archive_order":{"model":"SwinMTL","paper_title":"SwinMTL: A Shared Architecture for Simultaneous Depth Estimation and Semantic Segmentation from Monocular Camera Images","paper_url":"/paper/swinmtl-a-shared-architecture-for","paper_date":"2024-03-15","arxiv_id":"2403.10662","code_links":[{"title":"pardistaghavi/swinmtl","url":"https://github.com/pardistaghavi/swinmtl"}],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-sun-rgbd","slug":"monocular-depth-estimation-on-sun-rgbd","dataset":"SUN-RGBD","dataset_url":"/dataset/sun-rgb-d","rows_in_archive":3,"metrics":["Delta < 1.25","Delta < 1.25^2","Delta < 1.25^3","RMSE","absolute relative error","log 10"],"first_row_in_archive_order":{"model":"RPSF","paper_title":"End-to-end Learning for Joint Depth and Image Reconstruction from Diffracted Rotation","paper_url":"/paper/end-to-end-learning-for-joint-depth-and-image","paper_date":"2022-04-14","arxiv_id":"2204.07076","code_links":[],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-va","slug":"monocular-depth-estimation-on-va","dataset":"VA (Virtual Apartment)","dataset_url":"/dataset/va","rows_in_archive":3,"metrics":["Root mean square error (RMSE)","Log root mean square error  (RMSE_log)","Mean average error (MAE) ","Absolute relative error (AbsRel)"],"first_row_in_archive_order":{"model":"DistDepth","paper_title":"Toward Practical Monocular Indoor Depth Estimation","paper_url":"/paper/toward-practical-self-supervised-monocular","paper_date":"2021-12-04","arxiv_id":"2112.02306","code_links":[{"title":"facebookresearch/DistDepth","url":"https://github.com/facebookresearch/DistDepth"},{"title":"cake-lab/Mobile-AR-Depth-Estimation","url":"https://github.com/cake-lab/Mobile-AR-Depth-Estimation"}],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-kitti-2","slug":"monocular-depth-estimation-on-kitti-2","dataset":"KITTI","dataset_url":"/dataset/kitti","rows_in_archive":2,"metrics":["absolute relative error"],"first_row_in_archive_order":{"model":"MonoViT","paper_title":"MonoViT: Self-Supervised Monocular Depth Estimation with a Vision Transformer","paper_url":"/paper/monovit-self-supervised-monocular-depth","paper_date":"2022-08-06","arxiv_id":"2208.03543","code_links":[{"title":"zxcqlf/monovit","url":"https://github.com/zxcqlf/monovit"}],"syntology":{"n":6,"n_ran":6,"n_unverified":0,"n_pointer_only":0}}},{"leaderboard":"/sota/monocular-depth-estimation-on-middlebury-2014","slug":"monocular-depth-estimation-on-middlebury-2014","dataset":"Middlebury 2014","dataset_url":"/dataset/middlebury-2014","rows_in_archive":2,"metrics":["ORD ","D3R","RMSE","δ1.25"],"first_row_in_archive_order":{"model":"Miangoleh et al. (MiDaS)","paper_title":"Boosting Monocular Depth Estimation Models to High-Resolution via Content-Adaptive Multi-Resolution Merging","paper_url":"/paper/boosting-monocular-depth-estimation-models-to","paper_date":"2021-05-28","arxiv_id":"2105.14021","code_links":[{"title":"compphoto/BoostingMonocularDepth","url":"https://github.com/compphoto/BoostingMonocularDepth"}],"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}}},{"leaderboard":"/sota/monocular-depth-estimation-on-cityscapes-3d","slug":"monocular-depth-estimation-on-cityscapes-3d","dataset":"Cityscapes 3D","dataset_url":"/dataset/cityscapes-3d","rows_in_archive":1,"metrics":["RMSE"],"first_row_in_archive_order":{"model":"TaskPrompter","paper_title":"Joint 2D-3D Multi-Task Learning on Cityscapes-3D: 3D Detection, Segmentation, and Depth Estimation","paper_url":"/paper/joint-2d-3d-multi-task-learning-on-cityscapes","paper_date":"2023-04-03","arxiv_id":"2304.00971","code_links":[{"title":"prismformore/multi-task-transformer","url":"https://github.com/prismformore/multi-task-transformer"}],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-diml-outdoor","slug":"monocular-depth-estimation-on-diml-outdoor","dataset":"DIML Outdoor","dataset_url":null,"rows_in_archive":1,"metrics":["Delta < 1.25","RMSE","absolute relative error"],"first_row_in_archive_order":{"model":"ScaleDepth-NK","paper_title":"ScaleDepth: Decomposing Metric Depth Estimation into Scale Prediction and Relative Depth Estimation","paper_url":"/paper/scaledepth-decomposing-metric-depth","paper_date":"2024-07-11","arxiv_id":"2407.08187","code_links":[{"title":"RuijieZhu94/mmdepth","url":"https://github.com/RuijieZhu94/mmdepth/blob/main/projects/ScaleDepth/README.md"}],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-diode-indoor","slug":"monocular-depth-estimation-on-diode-indoor","dataset":"DIODE Indoor","dataset_url":null,"rows_in_archive":1,"metrics":["Delta < 1.25","RMSE","absolute relative error"],"first_row_in_archive_order":{"model":"ScaleDepth-NK","paper_title":"ScaleDepth: Decomposing Metric Depth Estimation into Scale Prediction and Relative Depth Estimation","paper_url":"/paper/scaledepth-decomposing-metric-depth","paper_date":"2024-07-11","arxiv_id":"2407.08187","code_links":[{"title":"RuijieZhu94/mmdepth","url":"https://github.com/RuijieZhu94/mmdepth/blob/main/projects/ScaleDepth/README.md"}],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-diode-outdoor","slug":"monocular-depth-estimation-on-diode-outdoor","dataset":"DIODE Outdoor","dataset_url":null,"rows_in_archive":1,"metrics":["Delta < 1.25","RMSE","absolute relative error"],"first_row_in_archive_order":{"model":"ScaleDepth-NK","paper_title":"ScaleDepth: Decomposing Metric Depth Estimation into Scale Prediction and Relative Depth Estimation","paper_url":"/paper/scaledepth-decomposing-metric-depth","paper_date":"2024-07-11","arxiv_id":"2407.08187","code_links":[{"title":"RuijieZhu94/mmdepth","url":"https://github.com/RuijieZhu94/mmdepth/blob/main/projects/ScaleDepth/README.md"}],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-hypersim","slug":"monocular-depth-estimation-on-hypersim","dataset":"Hypersim","dataset_url":"/dataset/hypersim","rows_in_archive":1,"metrics":["Delta < 1.25","RMSE","absolute relative error"],"first_row_in_archive_order":{"model":"ScaleDepth-NK","paper_title":"ScaleDepth: Decomposing Metric Depth Estimation into Scale Prediction and Relative Depth Estimation","paper_url":"/paper/scaledepth-decomposing-metric-depth","paper_date":"2024-07-11","arxiv_id":"2407.08187","code_links":[{"title":"RuijieZhu94/mmdepth","url":"https://github.com/RuijieZhu94/mmdepth/blob/main/projects/ScaleDepth/README.md"}],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-kitti-object","slug":"monocular-depth-estimation-on-kitti-object","dataset":"KITTI Object Tracking Evaluation 2012","dataset_url":"/dataset/kitti","rows_in_archive":1,"metrics":["Abs Rel"],"first_row_in_archive_order":{"model":"PackNet-SfM","paper_title":"3D Packing for Self-Supervised Monocular Depth Estimation","paper_url":"/paper/packnet-sfm-3d-packing-for-self-supervised","paper_date":"2019-05-06","arxiv_id":"1905.02693","code_links":[{"title":"TRI-ML/packnet-sfm","url":"https://github.com/TRI-ML/packnet-sfm"},{"title":"TRI-ML/DDAD","url":"https://github.com/TRI-ML/DDAD"},{"title":"ToyotaResearchInstitute/packnet-sfm","url":"https://github.com/ToyotaResearchInstitute/packnet-sfm"},{"title":"sejong-rcv/2021.Paper.TransDSSL","url":"https://github.com/sejong-rcv/2021.Paper.TransDSSL"}],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-matterport3d","slug":"monocular-depth-estimation-on-matterport3d","dataset":"Matterport3D","dataset_url":"/dataset/matterport3d","rows_in_archive":1,"metrics":["Delta < 1.25","Delta < 1.25^2","Delta < 1.25^3","RMSE","absolute error","absolute relative error"],"first_row_in_archive_order":{"model":"NeWCRFs","paper_title":"NeW CRFs: Neural Window Fully-connected CRFs for Monocular Depth Estimation","paper_url":"/paper/new-crfs-neural-window-fully-connected-crfs-1","paper_date":"2022-03-03","arxiv_id":"2203.01502","code_links":[{"title":"aliyun/NeWCRFs","url":"https://github.com/aliyun/NeWCRFs"}],"syntology":{"n":10,"n_ran":5,"n_unverified":5,"n_pointer_only":10}}},{"leaderboard":"/sota/monocular-depth-estimation-on-uasol","slug":"monocular-depth-estimation-on-uasol","dataset":"UASOL","dataset_url":"/dataset/uasol","rows_in_archive":1,"metrics":["RMSE"],"first_row_in_archive_order":{"model":"FCRN-DepthPrediction from Iro Laina et al. (2016)","paper_title":"UASOL, a large-scale high-resolution outdoor stereo dataset","paper_url":"/paper/uasol-a-large-scale-high-resolution-outdoor","paper_date":"2019-08-29","arxiv_id":null,"code_links":[],"syntology":null}},{"leaderboard":"/sota/monocular-depth-estimation-on-virtual-kitti-2","slug":"monocular-depth-estimation-on-virtual-kitti-2","dataset":"Virtual KITTI 2","dataset_url":"/dataset/virtual-kitti-2","rows_in_archive":1,"metrics":["Delta < 1.25","RMSE","absolute relative error"],"first_row_in_archive_order":{"model":"ScaleDepth-NK","paper_title":"ScaleDepth: Decomposing Metric Depth Estimation into Scale Prediction and Relative Depth Estimation","paper_url":"/paper/scaledepth-decomposing-metric-depth","paper_date":"2024-07-11","arxiv_id":"2407.08187","code_links":[{"title":"RuijieZhu94/mmdepth","url":"https://github.com/RuijieZhu94/mmdepth/blob/main/projects/ScaleDepth/README.md"}],"syntology":null}},{"leaderboard":null,"slug":"monocular-depth-estimation-on-mix-6-1","dataset":"MIX-6","dataset_url":null,"rows_in_archive":0,"metrics":["Zero-shot transfer"],"first_row_in_archive_order":null}],"datasets":[{"url":"/dataset/cityscapes","name":"Cityscapes","full_name":"","num_papers_in_archive":3702},{"url":"/dataset/kitti","name":"KITTI","full_name":"","num_papers_in_archive":3661},{"url":"/dataset/nyuv2","name":"NYUv2","full_name":"NYU-Depth V2","num_papers_in_archive":986},{"url":"/dataset/sun-rgb-d","name":"SUN RGB-D","full_name":"SUN RGB-D","num_papers_in_archive":477},{"url":"/dataset/matterport3d","name":"Matterport3D","full_name":"","num_papers_in_archive":461},{"url":"/dataset/virtual-kitti","name":"Virtual KITTI","full_name":"","num_papers_in_archive":133},{"url":"/dataset/make3d","name":"Make3D","full_name":"","num_papers_in_archive":129},{"url":"/dataset/eth3d","name":"ETH3D","full_name":"","num_papers_in_archive":121},{"url":"/dataset/hypersim","name":"Hypersim","full_name":"","num_papers_in_archive":108},{"url":"/dataset/diode","name":"DIODE","full_name":"Dense Indoor and Outdoor Depth","num_papers_in_archive":86},{"url":"/dataset/ddad","name":"DDAD","full_name":"Dense Depth for Autonomous Driving","num_papers_in_archive":73},{"url":"/dataset/middlebury-2014","name":"Middlebury 2014","full_name":"Middlebury 2014","num_papers_in_archive":59},{"url":"/dataset/virtual-kitti-2","name":"Virtual KITTI 2","full_name":"","num_papers_in_archive":53},{"url":"/dataset/ibims-1","name":"IBims-1","full_name":"Independent benchmark images and matched scans v1","num_papers_in_archive":34},{"url":"/dataset/mannequinchallenge","name":"MannequinChallenge","full_name":"MannequinChallenge","num_papers_in_archive":28},{"url":"/dataset/redweb","name":"ReDWeb","full_name":"Relative Depth from Web","num_papers_in_archive":22},{"url":"/dataset/3d60","name":"3D60","full_name":"","num_papers_in_archive":18},{"url":"/dataset/3d-ken-burns","name":"3D Ken Burns","full_name":"","num_papers_in_archive":13},{"url":"/dataset/cityscapes-3d","name":"Cityscapes 3D","full_name":"","num_papers_in_archive":13},{"url":"/dataset/wsvd","name":"WSVD","full_name":"Web Stereo Video Dataset","num_papers_in_archive":13},{"url":"/dataset/holopix50k","name":"Holopix50k","full_name":"","num_papers_in_archive":12},{"url":"/dataset/hrwsi","name":"HRWSI","full_name":"High-Resolution Web Stereo Image","num_papers_in_archive":12},{"url":"/dataset/human4d","name":"HUMAN4D","full_name":"","num_papers_in_archive":9},{"url":"/dataset/dttd2","name":"DTTD-Mobile","full_name":"","num_papers_in_archive":8},{"url":"/dataset/mid-air-dataset","name":"Mid-Air Dataset","full_name":"","num_papers_in_archive":6},{"url":"/dataset/syns-patches","name":"SYNS-Patches","full_name":"","num_papers_in_archive":5},{"url":"/dataset/muad","name":"MUAD","full_name":"Multiple Uncertainties for Autonomous Driving","num_papers_in_archive":4},{"url":"/dataset/uasol","name":"UASOL","full_name":"A large-scale high-resolution outdoor stereo dataset","num_papers_in_archive":3},{"url":"/dataset/va","name":"VA (Virtual Apartment)","full_name":"","num_papers_in_archive":3},{"url":"/dataset/vbr","name":"VBR","full_name":"VBR: A Vision Benchmark in Rome","num_papers_in_archive":3},{"url":"/dataset/infraparis","name":"InfraParis","full_name":"","num_papers_in_archive":1},{"url":"/dataset/polarimetric-imaging-for-perception","name":"Polarimetric Imaging for Perception","full_name":"","num_papers_in_archive":1},{"url":"/dataset/scared-c","name":"SCARED-C","full_name":"SCARED-Corrupted","num_papers_in_archive":1}],"subtasks":[],"parent_tasks":[{"url":"/task/depth-estimation","name":"Depth Estimation"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":430,"tagged_in_all":876,"items":[{"url":"/paper/high-quality-monocular-depth-estimation-via","title":"High Quality Monocular Depth Estimation via Transfer Learning","date":"2018-12-31","arxiv_id":"1812.11941","repositories_listed":45,"syntology":{"n":23,"n_ran":5,"n_unverified":18,"n_pointer_only":3}},{"url":"/paper/dinov2-learning-robust-visual-features","title":"DINOv2: Learning Robust Visual Features without Supervision","date":"2023-04-14","arxiv_id":"2304.07193","repositories_listed":26,"syntology":{"n":46,"n_ran":21,"n_unverified":25,"n_pointer_only":12}},{"url":"/paper/deeper-depth-prediction-with-fully","title":"Deeper Depth Prediction with Fully Convolutional Residual Networks","date":"2016-06-01","arxiv_id":"1606.00373","repositories_listed":18,"syntology":{"n":5,"n_ran":0,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/towards-robust-monocular-depth-estimation","title":"Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer","date":"2019-07-02","arxiv_id":"1907.01341","repositories_listed":16,"syntology":{"n":17,"n_ran":6,"n_unverified":11,"n_pointer_only":0}},{"url":"/paper/unsupervised-monocular-depth-estimation-with","title":"Unsupervised Monocular Depth Estimation with Left-Right Consistency","date":"2016-09-13","arxiv_id":"1609.03677","repositories_listed":16,"syntology":{"n":12,"n_ran":3,"n_unverified":9,"n_pointer_only":5}},{"url":"/paper/vision-transformers-for-dense-prediction","title":"Vision Transformers for Dense Prediction","date":"2021-03-24","arxiv_id":"2103.13413","repositories_listed":15,"syntology":{"n":116,"n_ran":54,"n_unverified":62,"n_pointer_only":15}},{"url":"/paper/digging-into-self-supervised-monocular-depth","title":"Digging Into Self-Supervised Monocular Depth Estimation","date":"2018-06-04","arxiv_id":"1806.01260","repositories_listed":15,"syntology":{"n":24,"n_ran":17,"n_unverified":7,"n_pointer_only":6}},{"url":"/paper/from-big-to-small-multi-scale-local-planar","title":"From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation","date":"2019-07-24","arxiv_id":"1907.10326","repositories_listed":14,"syntology":{"n":15,"n_ran":5,"n_unverified":10,"n_pointer_only":4}},{"url":"/paper/adabins-depth-estimation-using-adaptive-bins","title":"AdaBins: Depth Estimation using Adaptive Bins","date":"2020-11-28","arxiv_id":"2011.14141","repositories_listed":11,"syntology":null},{"url":"/paper/depth-prediction-without-the-sensors","title":"Depth Prediction Without the Sensors: Leveraging Structure for Unsupervised Learning from Monocular Videos","date":"2018-11-15","arxiv_id":"1811.06152","repositories_listed":11,"syntology":{"n":3,"n_ran":1,"n_unverified":2,"n_pointer_only":3}},{"url":"/paper/depth-map-prediction-from-a-single-image","title":"Depth Map Prediction from a Single Image using a Multi-Scale Deep Network","date":"2014-06-09","arxiv_id":"1406.2283","repositories_listed":10,"syntology":{"n":4,"n_ran":1,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/depth-anything-unleashing-the-power-of-large","title":"Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data","date":"2024-01-19","arxiv_id":"2401.10891","repositories_listed":7,"syntology":{"n":11,"n_ran":3,"n_unverified":8,"n_pointer_only":0}},{"url":"/paper/zoedepth-zero-shot-transfer-by-combining","title":"ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth","date":"2023-02-23","arxiv_id":"2302.12288","repositories_listed":6,"syntology":{"n":16,"n_ran":10,"n_unverified":6,"n_pointer_only":1}},{"url":"/paper/index-network","title":"Index Network","date":"2019-08-11","arxiv_id":"1908.09895","repositories_listed":6,"syntology":null},{"url":"/paper/unsupervised-monocular-depth-learning-in","title":"Unsupervised Monocular Depth Learning in Dynamic Scenes","date":"2020-10-30","arxiv_id":"2010.16404","repositories_listed":5,"syntology":{"n":14,"n_ran":0,"n_unverified":14,"n_pointer_only":0}},{"url":"/paper/deep-ordinal-regression-network-for-monocular","title":"Deep Ordinal Regression Network for Monocular Depth Estimation","date":"2018-06-06","arxiv_id":"1806.02446","repositories_listed":5,"syntology":{"n":3,"n_ran":0,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/repurposing-diffusion-based-image-generators","title":"Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation","date":"2023-12-04","arxiv_id":"2312.02145","repositories_listed":4,"syntology":{"n":26,"n_ran":14,"n_unverified":12,"n_pointer_only":0}},{"url":"/paper/global-local-path-networks-for-monocular","title":"Global-Local Path Networks for Monocular Depth Estimation with Vertical CutDepth","date":"2022-01-19","arxiv_id":"2201.07436","repositories_listed":4,"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/s2r-depthnet-learning-a-generalizable-depth","title":"S2R-DepthNet: Learning a Generalizable Depth-specific Structural Representation","date":"2021-04-02","arxiv_id":"2104.00877","repositories_listed":4,"syntology":null},{"url":"/paper/towards-better-generalization-joint-depth","title":"Towards Better Generalization: Joint Depth-Pose Learning without PoseNet","date":"2020-04-03","arxiv_id":"2004.01314","repositories_listed":4,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/packnet-sfm-3d-packing-for-self-supervised","title":"3D Packing for Self-Supervised Monocular Depth Estimation","date":"2019-05-06","arxiv_id":"1905.02693","repositories_listed":4,"syntology":null},{"url":"/paper/depth-from-videos-in-the-wild-unsupervised","title":"Depth from Videos in the Wild: Unsupervised Monocular Depth Learning from Unknown Cameras","date":"2019-04-10","arxiv_id":"1904.04998","repositories_listed":4,"syntology":{"n":4,"n_ran":1,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/fast-neural-architecture-search-of-compact","title":"Fast Neural Architecture Search of Compact Semantic Segmentation Models via Auxiliary Cells","date":"2018-10-25","arxiv_id":"1810.10804","repositories_listed":4,"syntology":null},{"url":"/paper/real-time-joint-semantic-segmentation-and","title":"Real-Time Joint Semantic Segmentation and Depth Estimation Using Asymmetric Annotations","date":"2018-09-13","arxiv_id":"1809.04766","repositories_listed":4,"syntology":{"n":6,"n_ran":6,"n_unverified":0,"n_pointer_only":6}},{"url":"/paper/towards-real-time-unsupervised-monocular","title":"Towards real-time unsupervised monocular depth estimation on CPU","date":"2018-06-29","arxiv_id":"1806.11430","repositories_listed":4,"syntology":null},{"url":"/paper/revisiting-single-image-depth-estimation","title":"Revisiting Single Image Depth Estimation: Toward Higher Resolution Maps with Accurate Object Boundaries","date":"2018-03-23","arxiv_id":"1803.08673","repositories_listed":4,"syntology":null},{"url":"/paper/fast-robust-monocular-depth-estimation-for","title":"Fast Robust Monocular Depth Estimation for Obstacle Detection with Fully Convolutional Networks","date":"2016-07-21","arxiv_id":"1607.06349","repositories_listed":4,"syntology":null},{"url":"/paper/predicting-depth-surface-normals-and-semantic","title":"Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-Scale Convolutional Architecture","date":"2014-11-18","arxiv_id":"1411.4734","repositories_listed":4,"syntology":{"n":3,"n_ran":0,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/on-the-robustness-of-language-guidance-for","title":"On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation","date":"2024-04-12","arxiv_id":"2404.08540","repositories_listed":3,"syntology":null},{"url":"/paper/unidepth-universal-monocular-metric-depth","title":"UniDepth: Universal Monocular Metric Depth Estimation","date":"2024-03-27","arxiv_id":"2403.18913","repositories_listed":3,"syntology":{"n":26,"n_ran":15,"n_unverified":11,"n_pointer_only":25}}],"syntology_records":21,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}