{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/monocular-depth-estimation-by-learning-from","title":"Monocular Depth Estimation by Learning from Heterogeneous Datasets","arxiv_id":"1803.08018","date":"2018-03-21","proceeding":null,"authors":["Akhil Gurram","Onay Urfalioglu","Ibrahim Halfaoui","Fahd Bouzaraa","Antonio M. Lopez"],"abstract":"Depth estimation provides essential information to perform autonomous driving\nand driver assistance. Especially, Monocular Depth Estimation is interesting\nfrom a practical point of view, since using a single camera is cheaper than\nmany other options and avoids the need for continuous calibration strategies as\nrequired by stereo-vision approaches. State-of-the-art methods for Monocular\nDepth Estimation are based on Convolutional Neural Networks (CNNs). A promising\nline of work consists of introducing additional semantic information about the\ntraffic scene when training CNNs for depth estimation. In practice, this means\nthat the depth data used for CNN training is complemented with images having\npixel-wise semantic labels, which usually are difficult to annotate (e.g.\ncrowded urban images). Moreover, so far it is common practice to assume that\nthe same raw training data is associated with both types of ground truth, i.e.,\ndepth and semantic labels. The main contribution of this paper is to show that\nthis hard constraint can be circumvented, i.e., that we can train CNNs for\ndepth estimation by leveraging the depth and semantic information coming from\nheterogeneous datasets. In order to illustrate the benefits of our approach, we\ncombine KITTI depth and Cityscapes semantic segmentation datasets,\noutperforming state-of-the-art results on Monocular Depth Estimation.","url_abs":"http://arxiv.org/abs/1803.08018v2","url_pdf":"http://arxiv.org/pdf/1803.08018v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"monocular-depth-estimation","task_name":"Monocular Depth Estimation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/monocular-depth-estimation-on-kitti-eigen","task":"Monocular Depth Estimation","dataset":"KITTI Eigen split","model":"CFA","rank_in_archive_order":55,"of":79,"metrics":{"absolute relative error":"0.096"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}