{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-comprehensive-representation","title":"Towards Comprehensive Representation Enhancement in Semantics-guided Self-supervised Monocular Depth Estimation","arxiv_id":null,"date":"2022-10-23","proceeding":"ECCV 2022 10","authors":["Jingyuan Ma","Xiangyu Lei","Nan Liu","Xian Zhao","ShiLiang Pu"],"abstract":"Semantics-guided self-supervised monocular depth estimation has been widely researched, owing to the strong cross-task correlation of depth and semantics. However, since depth estimation and\r\nsemantic segmentation are fundamentally two types of tasks: one is regression while the other is classification, the distribution of depth feature\r\nand semantic feature are naturally different. Previous works that leverage semantic information in depth estimation mostly neglect such representational discrimination, which leads to insufficient representation\r\nenhancement of depth feature. In this work, we propose an attentionbased module to enhance task-specific feature by addressing their feature\r\nuniqueness within instances. Additionally, we propose a metric learning based approach to accomplish comprehensive enhancement on depth\r\nfeature by creating a separation between instances in feature space. Extensive experiments and analysis demonstrate the effectiveness of our\r\nproposed method. In the end, our method achieves the state-of-the-art\r\nperformance on KITTI dataset.","url_abs":"https://link.springer.com/chapter/10.1007/978-3-031-19769-7_18","url_pdf":"https://www.ecva.net/papers/eccv_2022/papers_ECCV/papers/136610299.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"metric-learning","task_name":"Metric Learning"},{"task_slug":"monocular-depth-estimation","task_name":"Monocular Depth Estimation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"unsupervised-monocular-depth-estimation","task_name":"Unsupervised Monocular Depth Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/monocular-depth-estimation-on-kitti-eigen-1","task":"Monocular Depth Estimation","dataset":"KITTI Eigen split unsupervised","model":"CREMono(M + 1024x320 + Res50)","rank_in_archive_order":24,"of":55,"metrics":{"Delta < 1.25":"0.902","Delta < 1.25^2":" 0.969","Delta < 1.25^3":" 0.986","RMSE":"4.165","RMSE log":"0.171","Sq Rel":"0.624","absolute relative error":"0.099"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}