{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/activestereonet-end-to-end-self-supervised","title":"ActiveStereoNet: End-to-End Self-Supervised Learning for Active Stereo Systems","arxiv_id":"1807.06009","date":"2018-07-16","proceeding":"ECCV 2018 9","authors":["Yinda Zhang","Sameh Khamis","Christoph Rhemann","Julien Valentin","Adarsh Kowdle","Vladimir Tankovich","Michael Schoenberg","Shahram Izadi","Thomas Funkhouser","Sean Fanello"],"abstract":"In this paper we present ActiveStereoNet, the first deep learning solution\nfor active stereo systems. Due to the lack of ground truth, our method is fully\nself-supervised, yet it produces precise depth with a subpixel precision of\n$1/30th$ of a pixel; it does not suffer from the common over-smoothing issues;\nit preserves the edges; and it explicitly handles occlusions. We introduce a\nnovel reconstruction loss that is more robust to noise and texture-less\npatches, and is invariant to illumination changes. The proposed loss is\noptimized using a window-based cost aggregation with an adaptive support weight\nscheme. This cost aggregation is edge-preserving and smooths the loss function,\nwhich is key to allow the network to reach compelling results. Finally we show\nhow the task of predicting invalid regions, such as occlusions, can be trained\nend-to-end without ground-truth. This component is crucial to reduce blur and\nparticularly improves predictions along depth discontinuities. Extensive\nquantitatively and qualitatively evaluations on real and synthetic data\ndemonstrate state of the art results in many challenging scenes.","url_abs":"http://arxiv.org/abs/1807.06009v1","url_pdf":"http://arxiv.org/pdf/1807.06009v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"activestereonet-end-to-end-self-supervised","repo_url":"https://github.com/meteorshowers/X-StereoLab","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"self-supervised-learning","task_name":"Self-Supervised Learning"},{"task_slug":"stereo-disparity-estimation","task_name":"Stereo Disparity Estimation"},{"task_slug":"stereo-matching-1","task_name":"Stereo Matching"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1807.06009","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}