{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-monocular-depth-estimation-infusing","title":"Learning monocular depth estimation infusing traditional stereo knowledge","arxiv_id":"1904.04144","date":"2019-04-08","proceeding":"CVPR 2019 6","authors":["Fabio Tosi","Filippo Aleotti","Matteo Poggi","Stefano Mattoccia"],"abstract":"Depth estimation from a single image represents a fascinating, yet\nchallenging problem with countless applications. Recent works proved that this\ntask could be learned without direct supervision from ground truth labels\nleveraging image synthesis on sequences or stereo pairs. Focusing on this\nsecond case, in this paper we leverage stereo matching in order to improve\nmonocular depth estimation. To this aim we propose monoResMatch, a novel deep\narchitecture designed to infer depth from a single input image by synthesizing\nfeatures from a different point of view, horizontally aligned with the input\nimage, performing stereo matching between the two cues. In contrast to previous\nworks sharing this rationale, our network is the first trained end-to-end from\nscratch. Moreover, we show how obtaining proxy ground truth annotation through\ntraditional stereo algorithms, such as Semi-Global Matching, enables more\naccurate monocular depth estimation still countering the need for expensive\ndepth labels by keeping a self-supervised approach. Exhaustive experimental\nresults prove how the synergy between i) the proposed monoResMatch architecture\nand ii) proxy-supervision attains state-of-the-art for self-supervised\nmonocular depth estimation. The code is publicly available at\nhttps://github.com/fabiotosi92/monoResMatch-Tensorflow.","url_abs":"http://arxiv.org/abs/1904.04144v1","url_pdf":"http://arxiv.org/pdf/1904.04144v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-monocular-depth-estimation-infusing","repo_url":"https://github.com/fabiotosi92/monoResMatch-Tensorflow","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"monocular-depth-estimation","task_name":"Monocular Depth Estimation"},{"task_slug":"stereo-matching-1","task_name":"Stereo Matching"},{"task_slug":"stereo-matching","task_name":"Stereo Matching Hand"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/monocular-depth-estimation-on-kitti-eigen","task":"Monocular Depth Estimation","dataset":"KITTI Eigen split","model":"monoResMatch","rank_in_archive_order":53,"of":79,"metrics":{"absolute relative error":"0.096"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.04144","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}