{"url":"/task/depth-estimation","name":"Depth Estimation","slug":"depth-estimation","description_markdown":"**Depth Estimation** is the task of measuring the distance of each pixel relative to the camera. Depth is extracted from either monocular (single) or stereo (multiple views of a scene) images. Traditional methods use multi-view geometry to find the relationship between the images. Newer methods can directly estimate depth by minimizing the regression loss, or by learning to generate a novel view from a sequence. The most popular benchmarks are KITTI and NYUv2. Models are typically evaluated according to a RMS metric.\r\n\r\n<span class=\"description-source\">Source: [DIODE: A Dense Indoor and Outdoor DEpth Dataset ](https://arxiv.org/abs/1908.00463)</span>","categories":[{"name":"Computer Vision","url":"/area/computer-vision"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":2454,"papers_with_code":1029,"benchmarks":14,"benchmark_tables_in_archive":14,"benchmark_tables_shown":14,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":77,"subtasks":10,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/depth-estimation-on-stanford2d3d-panoramic","slug":"depth-estimation-on-stanford2d3d-panoramic","dataset":"Stanford2D3D Panoramic","dataset_url":"/dataset/2d-3d-s","rows_in_archive":18,"metrics":["RMSE","absolute relative error"],"first_row_in_archive_order":{"model":"HiMODE","paper_title":"HiMODE: A Hybrid Monocular Omnidirectional Depth Estimation Model","paper_url":"/paper/himode-a-hybrid-monocular-omnidirectional","paper_date":"2022-04-11","arxiv_id":"2204.05007","code_links":[],"syntology":null}},{"leaderboard":"/sota/depth-estimation-on-nyu-depth-v2","slug":"depth-estimation-on-nyu-depth-v2","dataset":"NYU-Depth V2","dataset_url":"/dataset/nyuv2","rows_in_archive":17,"metrics":["RMS","RMSE","mAP"],"first_row_in_archive_order":{"model":"EVP","paper_title":"EVP: Enhanced Visual Perception using Inverse Multi-Attentive Feature Refinement and Regularized Image-Text Alignment","paper_url":"/paper/evp-enhanced-visual-perception-using-inverse","paper_date":"2023-12-13","arxiv_id":"2312.08548","code_links":[{"title":"lavreniuk/evp","url":"https://github.com/lavreniuk/evp"}],"syntology":null}},{"leaderboard":"/sota/depth-estimation-on-dcm","slug":"depth-estimation-on-dcm","dataset":"DCM","dataset_url":"/dataset/dcm","rows_in_archive":3,"metrics":["Abs Rel","RMSE","RMSE log","Sq Rel"],"first_row_in_archive_order":{"model":"Bhattacharjee et al.","paper_title":"Estimating Image Depth in the Comics Domain","paper_url":"/paper/estimating-image-depth-in-the-comics-domain","paper_date":"2021-10-07","arxiv_id":"2110.03575","code_links":[{"title":"IVRL/ComicsDepth","url":"https://github.com/IVRL/ComicsDepth"}],"syntology":null}},{"leaderboard":"/sota/depth-estimation-on-ebdtheque","slug":"depth-estimation-on-ebdtheque","dataset":"eBDtheque","dataset_url":"/dataset/ebdtheque","rows_in_archive":3,"metrics":["Abs Rel","RMSE","RMSE log","Sq Rel"],"first_row_in_archive_order":{"model":"Bhattacharjee et al.","paper_title":"Estimating Image Depth in the Comics Domain","paper_url":"/paper/estimating-image-depth-in-the-comics-domain","paper_date":"2021-10-07","arxiv_id":"2110.03575","code_links":[{"title":"IVRL/ComicsDepth","url":"https://github.com/IVRL/ComicsDepth"}],"syntology":null}},{"leaderboard":"/sota/depth-estimation-on-scannetv2","slug":"depth-estimation-on-scannetv2","dataset":"ScanNetV2","dataset_url":"/dataset/scannet","rows_in_archive":3,"metrics":["absolute relative error","Delta < 1.25"],"first_row_in_archive_order":{"model":"Distill Any Depth","paper_title":"Distill Any Depth: Distillation Creates a Stronger Monocular Depth Estimator","paper_url":"/paper/distill-any-depth-distillation-creates-a","paper_date":"2025-02-26","arxiv_id":"2502.19204","code_links":[{"title":"Westlake-AGI-Lab/Distill-Any-Depth","url":"https://github.com/Westlake-AGI-Lab/Distill-Any-Depth"}],"syntology":null}},{"leaderboard":"/sota/depth-estimation-on-cityscapes-test","slug":"depth-estimation-on-cityscapes-test","dataset":"Cityscapes test","dataset_url":"/dataset/cityscapes","rows_in_archive":2,"metrics":["RMSE"],"first_row_in_archive_order":{"model":"SwinMTL","paper_title":"SwinMTL: A Shared Architecture for Simultaneous Depth Estimation and Semantic Segmentation from Monocular Camera Images","paper_url":"/paper/swinmtl-a-shared-architecture-for","paper_date":"2024-03-15","arxiv_id":"2403.10662","code_links":[{"title":"pardistaghavi/swinmtl","url":"https://github.com/pardistaghavi/swinmtl"}],"syntology":null}},{"leaderboard":"/sota/depth-estimation-on-diode","slug":"depth-estimation-on-diode","dataset":"DIODE","dataset_url":"/dataset/diode","rows_in_archive":2,"metrics":["Delta < 1.25","Delta < 1.25^2","Delta < 1.25^3"],"first_row_in_archive_order":{"model":"AIP-Brown","paper_title":"AI Playground: Unreal Engine-based Data Ablation Tool for Deep Learning","paper_url":"/paper/ai-playground-unreal-engine-based-data","paper_date":"2020-07-13","arxiv_id":"2007.06153","code_links":[{"title":"MMehdiMousavi/AIP","url":"https://github.com/MMehdiMousavi/AIP"}],"syntology":null}},{"leaderboard":"/sota/depth-estimation-on-kitti-2015","slug":"depth-estimation-on-kitti-2015","dataset":"KITTI 2015","dataset_url":"/dataset/kitti","rows_in_archive":2,"metrics":["Absolute relative error (AbsRel)","Sq Rel","RMSE"],"first_row_in_archive_order":{"model":"H-Net (Ours) Full Eigen","paper_title":"H-Net: Unsupervised Attention-based Stereo Depth Estimation Leveraging Epipolar Geometry","paper_url":"/paper/h-net-unsupervised-attention-based-stereo","paper_date":"2021-04-22","arxiv_id":"2104.11288","code_links":[],"syntology":null}},{"leaderboard":"/sota/depth-estimation-on-mars-dtm-estimation","slug":"depth-estimation-on-mars-dtm-estimation","dataset":"Mars DTM Estimation","dataset_url":"/dataset/mars-dtm-estimation","rows_in_archive":2,"metrics":["Delta < 1.25","Delta < 1.25^2","Delta < 1.25^3","Average PSNR","mean absolute error","RMSE"],"first_row_in_archive_order":{"model":"GLPDepth","paper_title":"An Adversarial Generative Network Designed for High-Resolution Monocular Depth Estimation from 2D HiRISE Images of Mars","paper_url":"/paper/an-adversarial-generative-network-designed","paper_date":"2022-08-15","arxiv_id":null,"code_links":[{"title":"riccardo2468/srdinet","url":"https://gitlab.com/riccardo2468/srdinet"}],"syntology":null}},{"leaderboard":"/sota/depth-estimation-on-scannet","slug":"depth-estimation-on-scannet","dataset":"ScanNet","dataset_url":"/dataset/scannet","rows_in_archive":2,"metrics":["RMSE","absolute relative error"],"first_row_in_archive_order":{"model":"Atlas (plain)","paper_title":"Atlas: End-to-End 3D Scene Reconstruction from Posed Images","paper_url":"/paper/atlas-end-to-end-3d-scene-reconstruction-from","paper_date":"2020-03-23","arxiv_id":"2003.10432","code_links":[{"title":"magicleap/Atlas","url":"https://github.com/magicleap/Atlas"}],"syntology":{"n":12,"n_ran":1,"n_unverified":11,"n_pointer_only":0}}},{"leaderboard":"/sota/depth-estimation-on-4d-light-field-dataset","slug":"depth-estimation-on-4d-light-field-dataset","dataset":"4D Light Field Dataset","dataset_url":"/dataset/4d-light-field-dataset","rows_in_archive":1,"metrics":["BadPix(0.01)","BadPix(0.03)","BadPix(0.07)","MSE "],"first_row_in_archive_order":{"model":"LFattNet","paper_title":"Attention-based View Selection Networks for Light-field Disparity Estimation","paper_url":"/paper/attention-based-view-selection-networks-for","paper_date":"2020-02-07","arxiv_id":null,"code_links":[{"title":"LIAGM/LFattNet","url":"https://github.com/LIAGM/LFattNet"}],"syntology":null}},{"leaderboard":"/sota/depth-estimation-on-kitti-eigen-split-3","slug":"depth-estimation-on-kitti-eigen-split-3","dataset":"KITTI Eigen split","dataset_url":"/dataset/kitti","rows_in_archive":1,"metrics":["Number of parameters (M)"],"first_row_in_archive_order":{"model":"LightDepth","paper_title":"LightDepth: A Resource Efficient Depth Estimation Approach for Dealing with Ground Truth Sparsity via Curriculum Learning","paper_url":"/paper/lightdepth-a-resource-efficient-depth","paper_date":"2022-11-16","arxiv_id":"2211.08608","code_links":[{"title":"fatemehkarimii/lightdepth","url":"https://github.com/fatemehkarimii/lightdepth"}],"syntology":null}},{"leaderboard":"/sota/depth-estimation-on-matterport3d","slug":"depth-estimation-on-matterport3d","dataset":"Matterport3D","dataset_url":"/dataset/matterport3d","rows_in_archive":1,"metrics":["Abs Rel"],"first_row_in_archive_order":{"model":"UniFuse","paper_title":"UniFuse: Unidirectional Fusion for 360$^{\\circ}$ Panorama Depth Estimation","paper_url":"/paper/unifuse-unidirectional-fusion-for-360-circ","paper_date":"2021-02-06","arxiv_id":"2102.03550","code_links":[{"title":"alibaba/UniFuse-Unidirectional-Fusion","url":"https://github.com/alibaba/UniFuse-Unidirectional-Fusion"}],"syntology":null}},{"leaderboard":"/sota/depth-estimation-on-taskonomy","slug":"depth-estimation-on-taskonomy","dataset":"Taskonomy","dataset_url":"/dataset/taskonomy","rows_in_archive":1,"metrics":["L1 error"],"first_row_in_archive_order":{"model":"X-TC (Cross-Task Consistency)","paper_title":"Robust Learning Through Cross-Task Consistency","paper_url":"/paper/robust-learning-through-cross-task","paper_date":"2020-06-01","arxiv_id":null,"code_links":[{"title":"EPFL-VILAB/XTConsistency","url":"https://github.com/EPFL-VILAB/XTConsistency"}],"syntology":null}}],"datasets":[{"url":"/dataset/cityscapes","name":"Cityscapes","full_name":"","num_papers_in_archive":3702},{"url":"/dataset/kitti","name":"KITTI","full_name":"","num_papers_in_archive":3661},{"url":"/dataset/scannet","name":"ScanNet","full_name":"","num_papers_in_archive":1595},{"url":"/dataset/nyuv2","name":"NYUv2","full_name":"NYU-Depth V2","num_papers_in_archive":986},{"url":"/dataset/matterport3d","name":"Matterport3D","full_name":"","num_papers_in_archive":461},{"url":"/dataset/tum-rgb-d","name":"TUM RGB-D","full_name":"TUM RGB-D","num_papers_in_archive":235},{"url":"/dataset/middlebury","name":"Middlebury","full_name":"Middlebury Stereo","num_papers_in_archive":223},{"url":"/dataset/suncg","name":"SUNCG","full_name":"SUNCG","num_papers_in_archive":186},{"url":"/dataset/megadepth","name":"MegaDepth","full_name":"","num_papers_in_archive":152},{"url":"/dataset/2d-3d-s","name":"2D-3D-S","full_name":"2D-3D-Semantic","num_papers_in_archive":147},{"url":"/dataset/taskonomy","name":"Taskonomy","full_name":"","num_papers_in_archive":147},{"url":"/dataset/virtual-kitti","name":"Virtual KITTI","full_name":"","num_papers_in_archive":133},{"url":"/dataset/make3d","name":"Make3D","full_name":"","num_papers_in_archive":129},{"url":"/dataset/sun3d","name":"SUN3D","full_name":"SUN3D","num_papers_in_archive":126},{"url":"/dataset/eth3d","name":"ETH3D","full_name":"","num_papers_in_archive":121},{"url":"/dataset/hypersim","name":"Hypersim","full_name":"","num_papers_in_archive":108},{"url":"/dataset/diode","name":"DIODE","full_name":"Dense Indoor and Outdoor Depth","num_papers_in_archive":86},{"url":"/dataset/ddad","name":"DDAD","full_name":"Dense Depth for Autonomous Driving","num_papers_in_archive":73},{"url":"/dataset/middlebury-2014","name":"Middlebury 2014","full_name":"Middlebury 2014","num_papers_in_archive":59},{"url":"/dataset/virtual-kitti-2","name":"Virtual KITTI 2","full_name":"","num_papers_in_archive":53},{"url":"/dataset/drivingstereo","name":"DrivingStereo","full_name":"","num_papers_in_archive":50},{"url":"/dataset/dense","name":"DENSE","full_name":"Depth Estimation oN Synthetic Events","num_papers_in_archive":49},{"url":"/dataset/mannequinchallenge","name":"MannequinChallenge","full_name":"MannequinChallenge","num_papers_in_archive":28},{"url":"/dataset/oasis","name":"OASIS","full_name":"Open Annotations of Single Image Surfaces","num_papers_in_archive":28},{"url":"/dataset/2d-3d-match-dataset","name":"2D-3D Match Dataset","full_name":"","num_papers_in_archive":27},{"url":"/dataset/void","name":"VOID","full_name":"Visual Odometry with Inertial and Depth","num_papers_in_archive":26},{"url":"/dataset/redweb","name":"ReDWeb","full_name":"Relative Depth from Web","num_papers_in_archive":22},{"url":"/dataset/3d60","name":"3D60","full_name":"","num_papers_in_archive":18},{"url":"/dataset/kitti-depth","name":"KITTI-Depth","full_name":null,"num_papers_in_archive":14},{"url":"/dataset/wsvd","name":"WSVD","full_name":"Web Stereo Video Dataset","num_papers_in_archive":13},{"url":"/dataset/holopix50k","name":"Holopix50k","full_name":"","num_papers_in_archive":12},{"url":"/dataset/hrwsi","name":"HRWSI","full_name":"High-Resolution Web Stereo Image","num_papers_in_archive":12},{"url":"/dataset/depth-in-the-wild","name":"Depth in the Wild","full_name":"","num_papers_in_archive":11},{"url":"/dataset/dynamic-replica","name":"Dynamic Replica","full_name":"","num_papers_in_archive":10},{"url":"/dataset/human4d","name":"HUMAN4D","full_name":"","num_papers_in_archive":9},{"url":"/dataset/dttd2","name":"DTTD-Mobile","full_name":"","num_papers_in_archive":8},{"url":"/dataset/tiktok-dataset","name":"TikTok Dataset","full_name":"Learning High Fidelity Depths of Dressed Humans  by Watching Social Media Dance Videos","num_papers_in_archive":8},{"url":"/dataset/stanford-light-field","name":"Stanford Light Field","full_name":"","num_papers_in_archive":7},{"url":"/dataset/durlar","name":"DurLAR","full_name":"A High-Fidelity 128-Channel LiDAR Dataset with Panoramic Ambient and Reflectivity Imagery","num_papers_in_archive":5},{"url":"/dataset/middlebury-2006","name":"Middlebury 2006","full_name":"Middlebury 2006","num_papers_in_archive":5},{"url":"/dataset/nerds-360","name":"NERDS 360","full_name":"NeRF for Reconstruction, Decomposition and Scene Synthesis of 360° outdoor scenes","num_papers_in_archive":5},{"url":"/dataset/serv-ct","name":"SERV-CT","full_name":"SERV-CT: A disparity dataset from CT for validation of endoscopic 3D reconstruction","num_papers_in_archive":5},{"url":"/dataset/syns-patches","name":"SYNS-Patches","full_name":"","num_papers_in_archive":5},{"url":"/dataset/3d-ken-burns-dataset","name":"3D Ken Burns Dataset","full_name":"","num_papers_in_archive":4},{"url":"/dataset/cocodoom","name":"CocoDoom","full_name":"","num_papers_in_archive":4},{"url":"/dataset/dcm","name":"DCM","full_name":"","num_papers_in_archive":4},{"url":"/dataset/ebdtheque","name":"eBDtheque","full_name":"","num_papers_in_archive":4},{"url":"/dataset/eden","name":"EDEN","full_name":"","num_papers_in_archive":4},{"url":"/dataset/endoslam","name":"EndoSLAM","full_name":"Endoscopic SLAM dataset","num_papers_in_archive":4},{"url":"/dataset/futurehouse","name":"FutureHouse","full_name":"","num_papers_in_archive":3},{"url":"/dataset/seasondepth","name":"SeasonDepth","full_name":"","num_papers_in_archive":3},{"url":"/dataset/uasol","name":"UASOL","full_name":"A large-scale high-resolution outdoor stereo dataset","num_papers_in_archive":3},{"url":"/dataset/4d-light-field-dataset","name":"4D Light Field Dataset","full_name":"","num_papers_in_archive":2},{"url":"/dataset/enrich","name":"ENRICH","full_name":"Multi-purposE dataset for beNchmaRking In Computer vision and pHotogrammetry","num_papers_in_archive":2},{"url":"/dataset/g-vue","name":"G-VUE","full_name":"General-purpose Visual Understanding Evaluation","num_papers_in_archive":2},{"url":"/dataset/hammer","name":"HAMMER","full_name":"","num_papers_in_archive":2},{"url":"/dataset/headcam","name":"Headcam","full_name":"","num_papers_in_archive":2},{"url":"/dataset/instaorder","name":"InstaOrder","full_name":"","num_papers_in_archive":2},{"url":"/dataset/mila-simulated-floods","name":"Mila Simulated Floods","full_name":"","num_papers_in_archive":2},{"url":"/dataset/pano3d","name":"Pano3D","full_name":"","num_papers_in_archive":2},{"url":"/dataset/supercaustics","name":"SuperCaustics","full_name":"","num_papers_in_archive":2},{"url":"/dataset/autonomous-driving-streaming-perception","name":"Autonomous-driving Streaming Perception Benchmarrk","full_name":"","num_papers_in_archive":1},{"url":"/dataset/carla2real","name":"CARLA2Real","full_name":"","num_papers_in_archive":1},{"url":"/dataset/coastal-inundation-maps-with-floodwater-depth","name":"Coastal Inundation Maps with Floodwater Depth Values","full_name":"Simulated Flood Inundation Maps of Abu Dhabi's Coast Under Different Shoreline Protection Scenarios","num_papers_in_archive":1},{"url":"/dataset/conslam","name":"ConSLAM","full_name":"Construction Dataset for SLAM","num_papers_in_archive":1},{"url":"/dataset/depth-from-couple-optical-differentiation","name":"Depth from Couple Optical Differentiation","full_name":"","num_papers_in_archive":1},{"url":"/dataset/dermsynth3d","name":"DermSynth3D","full_name":"3DBodyTex.DermSynth3D","num_papers_in_archive":1},{"url":"/dataset/indoor-and-outdoor-dfd-dataset","name":"Indoor and outdoor DFD dataset","full_name":"","num_papers_in_archive":1},{"url":"/dataset/inria-dlfd","name":"INRIA DLFD","full_name":"INRIA Dense Light Field","num_papers_in_archive":1},{"url":"/dataset/mars-dtm-estimation","name":"Mars DTM Estimation","full_name":"","num_papers_in_archive":1},{"url":"/dataset/minenav","name":"MineNav","full_name":"","num_papers_in_archive":1},{"url":"/dataset/scared-c","name":"SCARED-C","full_name":"SCARED-Corrupted","num_papers_in_archive":1},{"url":"/dataset/simbev","name":"SimBEV","full_name":"","num_papers_in_archive":1},{"url":"/dataset/transproteus","name":"TransProteus","full_name":"","num_papers_in_archive":1},{"url":"/dataset/vis-tir","name":"VIS-TIR","full_name":"","num_papers_in_archive":1},{"url":"/dataset/ms2-dataset-rgb-nir-thermal-images-lidar-gps","name":"Multi-Spectral Stereo Dataset  (RGB, NIR, thermal images, LiDAR, GPS/IMU)","full_name":"","num_papers_in_archive":0},{"url":"/dataset/pinet","name":"PINet","full_name":"","num_papers_in_archive":0}],"subtasks":[{"url":"/task/3d-depth-estimation","name":"3D Depth Estimation"},{"url":"/task/bathymetry-prediction","name":"Bathymetry prediction"},{"url":"/task/depth-aleatoric-uncertainty-estimation","name":"Depth Aleatoric Uncertainty Estimation"},{"url":"/task/depth-and-camera-motion","name":"Depth And Camera Motion"},{"url":"/task/depth-image-upsampling","name":"Depth Image Upsampling"},{"url":"/task/depth-map-super-resolution","name":"Depth Map Super-Resolution"},{"url":"/task/indoor-monocular-depth-estimation","name":"Indoor Monocular Depth Estimation"},{"url":"/task/monocular-depth-estimation","name":"Monocular Depth Estimation"},{"url":"/task/stereo-depth-estimation","name":"Stereo Depth Estimation"},{"url":"/task/stereo-lidar-fusion","name":"Stereo-LiDAR Fusion"}],"parent_tasks":[{"url":"/task/3d","name":"3D"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":1029,"tagged_in_all":2454,"items":[{"url":"/paper/high-quality-monocular-depth-estimation-via","title":"High Quality Monocular Depth Estimation via Transfer Learning","date":"2018-12-31","arxiv_id":"1812.11941","repositories_listed":45,"syntology":{"n":23,"n_ran":5,"n_unverified":18,"n_pointer_only":3}},{"url":"/paper/dinov2-learning-robust-visual-features","title":"DINOv2: Learning Robust Visual Features without Supervision","date":"2023-04-14","arxiv_id":"2304.07193","repositories_listed":26,"syntology":{"n":46,"n_ran":21,"n_unverified":25,"n_pointer_only":12}},{"url":"/paper/deeper-depth-prediction-with-fully","title":"Deeper Depth Prediction with Fully Convolutional Residual Networks","date":"2016-06-01","arxiv_id":"1606.00373","repositories_listed":18,"syntology":{"n":5,"n_ran":0,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/towards-robust-monocular-depth-estimation","title":"Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer","date":"2019-07-02","arxiv_id":"1907.01341","repositories_listed":16,"syntology":{"n":17,"n_ran":6,"n_unverified":11,"n_pointer_only":0}},{"url":"/paper/unsupervised-monocular-depth-estimation-with","title":"Unsupervised Monocular Depth Estimation with Left-Right Consistency","date":"2016-09-13","arxiv_id":"1609.03677","repositories_listed":16,"syntology":{"n":12,"n_ran":3,"n_unverified":9,"n_pointer_only":5}},{"url":"/paper/vision-transformers-for-dense-prediction","title":"Vision Transformers for Dense Prediction","date":"2021-03-24","arxiv_id":"2103.13413","repositories_listed":15,"syntology":{"n":116,"n_ran":54,"n_unverified":62,"n_pointer_only":15}},{"url":"/paper/digging-into-self-supervised-monocular-depth","title":"Digging Into Self-Supervised Monocular Depth Estimation","date":"2018-06-04","arxiv_id":"1806.01260","repositories_listed":15,"syntology":{"n":24,"n_ran":17,"n_unverified":7,"n_pointer_only":6}},{"url":"/paper/from-big-to-small-multi-scale-local-planar","title":"From Big to Small: Multi-Scale Local Planar Guidance for Monocular Depth Estimation","date":"2019-07-24","arxiv_id":"1907.10326","repositories_listed":14,"syntology":{"n":15,"n_ran":5,"n_unverified":10,"n_pointer_only":4}},{"url":"/paper/factorized-attention-self-attention-with","title":"Efficient Attention: Attention with Linear Complexities","date":"2018-12-04","arxiv_id":"1812.01243","repositories_listed":14,"syntology":{"n":8,"n_ran":7,"n_unverified":1,"n_pointer_only":1}},{"url":"/paper/adabins-depth-estimation-using-adaptive-bins","title":"AdaBins: Depth Estimation using Adaptive Bins","date":"2020-11-28","arxiv_id":"2011.14141","repositories_listed":11,"syntology":null},{"url":"/paper/depth-prediction-without-the-sensors","title":"Depth Prediction Without the Sensors: Leveraging Structure for Unsupervised Learning from Monocular Videos","date":"2018-11-15","arxiv_id":"1811.06152","repositories_listed":11,"syntology":{"n":3,"n_ran":1,"n_unverified":2,"n_pointer_only":3}},{"url":"/paper/what-uncertainties-do-we-need-in-bayesian","title":"What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?","date":"2017-03-15","arxiv_id":"1703.04977","repositories_listed":11,"syntology":{"n":5,"n_ran":4,"n_unverified":1,"n_pointer_only":4}},{"url":"/paper/depth-anything-unleashing-the-power-of-large","title":"Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data","date":"2024-01-19","arxiv_id":"2401.10891","repositories_listed":7,"syntology":{"n":11,"n_ran":3,"n_unverified":8,"n_pointer_only":0}},{"url":"/paper/self-supervised-learning-from-images-with-a","title":"Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture","date":"2023-01-19","arxiv_id":"2301.08243","repositories_listed":7,"syntology":{"n":14,"n_ran":5,"n_unverified":9,"n_pointer_only":13}},{"url":"/paper/multi-task-learning-as-multi-objective","title":"Multi-Task Learning as Multi-Objective Optimization","date":"2018-10-10","arxiv_id":"1810.04650","repositories_listed":7,"syntology":{"n":19,"n_ran":2,"n_unverified":17,"n_pointer_only":0}},{"url":"/paper/zoedepth-zero-shot-transfer-by-combining","title":"ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth","date":"2023-02-23","arxiv_id":"2302.12288","repositories_listed":6,"syntology":{"n":16,"n_ran":10,"n_unverified":6,"n_pointer_only":1}},{"url":"/paper/epp-mvsnet-epipolar-assembling-based-depth","title":"EPP-MVSNet: Epipolar-Assembling Based Depth Prediction for Multi-View Stereo","date":"2021-01-01","arxiv_id":null,"repositories_listed":6,"syntology":null},{"url":"/paper/index-network","title":"Index Network","date":"2019-08-11","arxiv_id":"1908.09895","repositories_listed":6,"syntology":null},{"url":"/paper/pyramid-stereo-matching-network","title":"Pyramid Stereo Matching Network","date":"2018-03-23","arxiv_id":"1803.08669","repositories_listed":6,"syntology":{"n":11,"n_ran":3,"n_unverified":8,"n_pointer_only":0}},{"url":"/paper/sparse-to-dense-depth-prediction-from-sparse","title":"Sparse-to-Dense: Depth Prediction from Sparse Depth Samples and a Single Image","date":"2017-09-21","arxiv_id":"1709.07492","repositories_listed":6,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/unsupervised-monocular-depth-learning-in","title":"Unsupervised Monocular Depth Learning in Dynamic Scenes","date":"2020-10-30","arxiv_id":"2010.16404","repositories_listed":5,"syntology":{"n":14,"n_ran":0,"n_unverified":14,"n_pointer_only":0}},{"url":"/paper/deep-ordinal-regression-network-for-monocular","title":"Deep Ordinal Regression Network for Monocular Depth Estimation","date":"2018-06-06","arxiv_id":"1806.02446","repositories_listed":5,"syntology":{"n":3,"n_ran":0,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/deep-depth-from-focus","title":"Deep Depth From Focus","date":"2017-04-04","arxiv_id":"1704.01085","repositories_listed":5,"syntology":null},{"url":"/paper/repurposing-diffusion-based-image-generators","title":"Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation","date":"2023-12-04","arxiv_id":"2312.02145","repositories_listed":4,"syntology":{"n":26,"n_ran":14,"n_unverified":12,"n_pointer_only":0}},{"url":"/paper/u-hrnet-delving-into-improving-semantic","title":"U-HRNet: Delving into Improving Semantic Representation of High Resolution Network for Dense Prediction","date":"2022-10-13","arxiv_id":"2210.07140","repositories_listed":4,"syntology":null},{"url":"/paper/expediting-large-scale-vision-transformer-for","title":"Expediting Large-Scale Vision Transformer for Dense Prediction without Fine-tuning","date":"2022-10-03","arxiv_id":"2210.01035","repositories_listed":4,"syntology":{"n":7,"n_ran":4,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/global-local-path-networks-for-monocular","title":"Global-Local Path Networks for Monocular Depth Estimation with Vertical CutDepth","date":"2022-01-19","arxiv_id":"2201.07436","repositories_listed":4,"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/mobilestereonet-towards-lightweight-deep","title":"MobileStereoNet: Towards Lightweight Deep Networks for Stereo Matching","date":"2021-08-22","arxiv_id":"2108.09770","repositories_listed":4,"syntology":{"n":8,"n_ran":1,"n_unverified":7,"n_pointer_only":0}},{"url":"/paper/s2r-depthnet-learning-a-generalizable-depth","title":"S2R-DepthNet: Learning a Generalizable Depth-specific Structural Representation","date":"2021-04-02","arxiv_id":"2104.00877","repositories_listed":4,"syntology":null},{"url":"/paper/towards-better-generalization-joint-depth","title":"Towards Better Generalization: Joint Depth-Pose Learning without PoseNet","date":"2020-04-03","arxiv_id":"2004.01314","repositories_listed":4,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":0}}],"syntology_records":24,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}