{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-optics-for-monocular-depth-estimation","title":"Deep Optics for Monocular Depth Estimation and 3D Object Detection","arxiv_id":"1904.08601","date":"2019-04-18","proceeding":"ICCV 2019 10","authors":["Julie Chang","Gordon Wetzstein"],"abstract":"Depth estimation and 3D object detection are critical for scene understanding\nbut remain challenging to perform with a single image due to the loss of 3D\ninformation during image capture. Recent models using deep neural networks have\nimproved monocular depth estimation performance, but there is still difficulty\nin predicting absolute depth and generalizing outside a standard dataset. Here\nwe introduce the paradigm of deep optics, i.e. end-to-end design of optics and\nimage processing, to the monocular depth estimation problem, using coded\ndefocus blur as an additional depth cue to be decoded by a neural network. We\nevaluate several optical coding strategies along with an end-to-end\noptimization scheme for depth estimation on three datasets, including NYU Depth\nv2 and KITTI. We find an optimized freeform lens design yields the best\nresults, but chromatic aberration from a singlet lens offers significantly\nimproved performance as well. We build a physical prototype and validate that\nchromatic aberrations improve depth estimation on real-world results. In\naddition, we train object detection networks on the KITTI dataset and show that\nthe lens optimized for depth estimation also results in improved 3D object\ndetection performance.","url_abs":"http://arxiv.org/abs/1904.08601v1","url_pdf":"http://arxiv.org/pdf/1904.08601v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-object-detection","task_name":"3D Object Detection"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"monocular-depth-estimation","task_name":"Monocular Depth Estimation"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/depth-estimation-on-nyu-depth-v2","task":"Depth Estimation","dataset":"NYU-Depth V2","model":"Optimized, freeform","rank_in_archive_order":11,"of":17,"metrics":{"RMS":"0.4325"},"uses_additional_data":false},{"leaderboard":"/sota/depth-estimation-on-nyu-depth-v2","task":"Depth Estimation","dataset":"NYU-Depth V2","model":"Freeform","rank_in_archive_order":12,"of":17,"metrics":{"RMS":"0.433"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.08601","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}