{"url":"/sota/monocular-depth-estimation-on-va","task":{"name":"Monocular Depth Estimation","url":"/task/monocular-depth-estimation","note":null},"dataset":{"name":"VA (Virtual Apartment)","url":"/dataset/va"},"category":"Computer Vision","categories":["Computer Vision"],"category_note":null,"description":"**Monocular Depth Estimation** is the task of estimating the depth value (distance relative to the camera) of each pixel given a single (monocular) RGB image. This challenging task is a key prerequisite for determining scene understanding for applications such as 3D scene reconstruction, autonomous driving, and AR. State-of-the-art methods usually fall into one of two categories: designing a complex network that is powerful enough to directly regress the depth map, or splitting the input into bins or windows to reduce computational complexity.  The most popular benchmarks are the KITTI and NYUv2 datasets. Models are typically evaluated using RMSE or absolute relative error. \r\n\r\n<span class=\"description-source\">Source: [Defocus Deblurring Using Dual-Pixel Data ](https://arxiv.org/abs/2005.00305)</span>","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["Root mean square error (RMSE)","Log root mean square error  (RMSE_log)","Mean average error (MAE) ","Absolute relative error (AbsRel)"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"Root mean square error (RMSE)":"lower","Log root mean square error  (RMSE_log)":"lower","Mean average error (MAE) ":"lower","Absolute relative error (AbsRel)":"lower"}},"counts":{"rows":3,"rows_with_code":3,"rows_with_paper_page":3,"rows_dated":3,"rows_using_additional_data":0},"rows":[{"rank_in_archive_order":1,"model":"DistDepth","metrics":{"Absolute relative error (AbsRel)":"0.175","Log root mean square error  (RMSE_log)":"0.213","Mean average error (MAE) ":"0.253","Root mean square error (RMSE)":"0.374"},"uses_additional_data":false,"paper_date":"2021-12-04","paper":"/paper/toward-practical-self-supervised-monocular","paper_url":"https://arxiv.org/abs/2112.02306v2","paper_title":"Toward Practical Monocular Indoor Depth Estimation","code":"https://github.com/facebookresearch/DistDepth","n_code_links":2,"syntology":null},{"rank_in_archive_order":2,"model":"Depth Hints","metrics":{"Absolute relative error (AbsRel)":"0.197","Log root mean square error  (RMSE_log)":"0.248","Mean average error (MAE) ":"0.291","Root mean square error (RMSE)":"0.427"},"uses_additional_data":false,"paper_date":"2019-09-19","paper":"/paper/self-supervised-monocular-depth-hints","paper_url":"https://arxiv.org/abs/1909.09051v1","paper_title":"Self-Supervised Monocular Depth Hints","code":"https://github.com/nianticlabs/depth-hints","n_code_links":1,"syntology":{"n_ran":2,"n_unverified":0,"n_samples":2,"n_pointer_only_licence":2}},{"rank_in_archive_order":3,"model":"MonoDepth2","metrics":{"Absolute relative error (AbsRel)":"0.203","Log root mean square error  (RMSE_log)":"0.251","Mean average error (MAE) ":"0.295","Root mean square error (RMSE)":"0.432"},"uses_additional_data":false,"paper_date":"2018-06-04","paper":"/paper/digging-into-self-supervised-monocular-depth","paper_url":"https://arxiv.org/abs/1806.01260v4","paper_title":"Digging Into Self-Supervised Monocular Depth Estimation","code":"https://github.com/nianticlabs/monodepth2","n_code_links":15,"syntology":{"n_ran":17,"n_unverified":7,"n_samples":24,"n_pointer_only_licence":6}}],"since_archive":{"present":false,"note":"No Syntology-extracted rows are published in this build."},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":2,"rows_with_any_sample_ran":2,"distinct_papers_with_graph_line":2,"distinct_papers_with_any_sample_ran":2,"samples_over_distinct_papers":{"n_ran":19,"n_unverified":7,"n_samples":26,"n_pointer_only_licence":8,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":19,"n_unverified":7,"n_samples":26,"n_pointer_only_licence":8,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}