{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/geometry-meets-semantics-for-semi-supervised","title":"Geometry meets semantics for semi-supervised monocular depth estimation","arxiv_id":"1810.04093","date":"2018-10-09","proceeding":null,"authors":["Pierluigi Zama Ramirez","Matteo Poggi","Fabio Tosi","Stefano Mattoccia","Luigi Di Stefano"],"abstract":"Depth estimation from a single image represents a very exciting challenge in\ncomputer vision. While other image-based depth sensing techniques leverage on\nthe geometry between different viewpoints (e.g., stereo or structure from\nmotion), the lack of these cues within a single image renders ill-posed the\nmonocular depth estimation task. For inference, state-of-the-art\nencoder-decoder architectures for monocular depth estimation rely on effective\nfeature representations learned at training time. For unsupervised training of\nthese models, geometry has been effectively exploited by suitable images\nwarping losses computed from views acquired by a stereo rig or a moving camera.\nIn this paper, we make a further step forward showing that learning semantic\ninformation from images enables to improve effectively monocular depth\nestimation as well. In particular, by leveraging on semantically labeled images\ntogether with unsupervised signals gained by geometry through an image warping\nloss, we propose a deep learning approach aimed at joint semantic segmentation\nand depth estimation. Our overall learning framework is semi-supervised, as we\ndeploy groundtruth data only in the semantic domain. At training time, our\nnetwork learns a common feature representation for both tasks and a novel\ncross-task loss function is proposed. The experimental findings show how,\njointly tackling depth prediction and semantic segmentation, allows to improve\ndepth estimation accuracy. In particular, on the KITTI dataset our network\noutperforms state-of-the-art methods for monocular depth estimation.","url_abs":"http://arxiv.org/abs/1810.04093v2","url_pdf":"http://arxiv.org/pdf/1810.04093v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"geometry-meets-semantics-for-semi-supervised","repo_url":"https://github.com/CVLAB-Unibo/Semantic-Mono-Depth","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"depth-prediction","task_name":"Depth Prediction"},{"task_slug":"monocular-depth-estimation","task_name":"Monocular Depth Estimation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1810.04093","atlas_url":"https://app.syntology.ai/?focus=1810.04093","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}