{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-auxiliary-learning-for-visual","title":"Deep Auxiliary Learning for Visual Localization and Odometry","arxiv_id":"1803.03642","date":"2018-03-09","proceeding":null,"authors":["Abhinav Valada","Noha Radwan","Wolfram Burgard"],"abstract":"Localization is an indispensable component of a robot's autonomy stack that\nenables it to determine where it is in the environment, essentially making it a\nprecursor for any action execution or planning. Although convolutional neural\nnetworks have shown promising results for visual localization, they are still\ngrossly outperformed by state-of-the-art local feature-based techniques. In\nthis work, we propose VLocNet, a new convolutional neural network architecture\nfor 6-DoF global pose regression and odometry estimation from consecutive\nmonocular images. Our multitask model incorporates hard parameter sharing, thus\nbeing compact and enabling real-time inference, in addition to being end-to-end\ntrainable. We propose a novel loss function that utilizes auxiliary learning to\nleverage relative pose information during training, thereby constraining the\nsearch space to obtain consistent pose estimates. We evaluate our proposed\nVLocNet on indoor as well as outdoor datasets and show that even our single\ntask model exceeds the performance of state-of-the-art deep architectures for\nglobal localization, while achieving competitive performance for visual\nodometry estimation. Furthermore, we present extensive experimental evaluations\nutilizing our proposed Geometric Consistency Loss that show the effectiveness\nof multitask learning and demonstrate that our model is the first deep learning\ntechnique to be on par with, and in some cases outperforms state-of-the-art\nSIFT-based approaches.","url_abs":"http://arxiv.org/abs/1803.03642v1","url_pdf":"http://arxiv.org/pdf/1803.03642v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-auxiliary-learning-for-visual","repo_url":"https://github.com/decayale/vlocnet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"auxiliary-learning","task_name":"Auxiliary Learning"},{"task_slug":"visual-localization","task_name":"Visual Localization"},{"task_slug":"visual-odometry","task_name":"Visual Odometry"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.03642","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}