{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-learning-of-depth-and-ego-motion-2","title":"Unsupervised Learning of Depth and Ego-Motion from Monocular Video Using 3D Geometric Constraints","arxiv_id":"1802.05522","date":"2018-02-15","proceeding":"CVPR 2018 6","authors":["Reza Mahjourian","Martin Wicke","Anelia Angelova"],"abstract":"We present a novel approach for unsupervised learning of depth and ego-motion\nfrom monocular video. Unsupervised learning removes the need for separate\nsupervisory signals (depth or ego-motion ground truth, or multi-view video).\nPrior work in unsupervised depth learning uses pixel-wise or gradient-based\nlosses, which only consider pixels in small local neighborhoods. Our main\ncontribution is to explicitly consider the inferred 3D geometry of the scene,\nenforcing consistency of the estimated 3D point clouds and ego-motion across\nconsecutive frames. This is a challenging task and is solved by a novel\n(approximate) backpropagation algorithm for aligning 3D structures.\n  We combine this novel 3D-based loss with 2D losses based on photometric\nquality of frame reconstructions using estimated depth and ego-motion from\nadjacent frames. We also incorporate validity masks to avoid penalizing areas\nin which no useful information exists.\n  We test our algorithm on the KITTI dataset and on a video dataset captured on\nan uncalibrated mobile phone camera. Our proposed approach consistently\nimproves depth estimates on both datasets, and outperforms the state-of-the-art\nfor both depth and ego-motion. Because we only require a simple video, learning\ndepth and ego-motion on large and varied datasets becomes possible. We\ndemonstrate this by training on the low quality uncalibrated video dataset and\nevaluating on KITTI, ranking among top performing prior methods which are\ntrained on KITTI itself.","url_abs":"http://arxiv.org/abs/1802.05522v2","url_pdf":"http://arxiv.org/pdf/1802.05522v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-learning-of-depth-and-ego-motion-2","repo_url":"https://github.com/tensorflow/models/tree/master/research/vid2depth","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"unsupervised-learning-of-depth-and-ego-motion-2","repo_url":"https://github.com/Shiaoming/vid2depth","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"3d-geometry","task_name":"3D geometry"},{"task_slug":"depth-and-camera-motion","task_name":"Depth And Camera Motion"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"monocular-depth-estimation","task_name":"Monocular Depth Estimation"},{"task_slug":"simultaneous-localization-and-mapping","task_name":"Simultaneous Localization and Mapping"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1802.05522","atlas_url":"https://app.syntology.ai/?focus=1802.05522","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}