{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/monocular-depth-estimation-with-hierarchical","title":"Monocular Depth Estimation with Hierarchical Fusion of Dilated CNNs and Soft-Weighted-Sum Inference","arxiv_id":"1708.02287","date":"2017-08-02","proceeding":null,"authors":["Bo Li","Yuchao Dai","Mingyi He"],"abstract":"Monocular depth estimation is a challenging task in complex compositions\ndepicting multiple objects of diverse scales. Albeit the recent great progress\nthanks to the deep convolutional neural networks (CNNs), the state-of-the-art\nmonocular depth estimation methods still fall short to handle such real-world\nchallenging scenarios. In this paper, we propose a deep end-to-end learning\nframework to tackle these challenges, which learns the direct mapping from a\ncolor image to the corresponding depth map. First, we represent monocular depth\nestimation as a multi-category dense labeling task by contrast to the\nregression based formulation. In this way, we could build upon the recent\nprogress in dense labeling such as semantic segmentation. Second, we fuse\ndifferent side-outputs from our front-end dilated convolutional neural network\nin a hierarchical way to exploit the multi-scale depth cues for depth\nestimation, which is critical to achieve scale-aware depth estimation. Third,\nwe propose to utilize soft-weighted-sum inference instead of the hard-max\ninference, transforming the discretized depth score to continuous depth value.\nThus, we reduce the influence of quantization error and improve the robustness\nof our method. Extensive experiments on the NYU Depth V2 and KITTI datasets\nshow the superiority of our method compared with current state-of-the-art\nmethods. Furthermore, experiments on the NYU V2 dataset reveal that our model\nis able to learn the probability distribution of depth.","url_abs":"http://arxiv.org/abs/1708.02287v1","url_pdf":"http://arxiv.org/pdf/1708.02287v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"monocular-depth-estimation-with-hierarchical","repo_url":"https://github.com/racinmat/depth-voxelmap-estimation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"monocular-depth-estimation","task_name":"Monocular Depth Estimation"},{"task_slug":"quantization","task_name":"Quantization"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1708.02287","atlas_url":"https://app.syntology.ai/?focus=1708.02287","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}