{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vlocnet-deep-multitask-learning-for-semantic","title":"VLocNet++: Deep Multitask Learning for Semantic Visual Localization and Odometry","arxiv_id":"1804.08366","date":"2018-04-23","proceeding":null,"authors":["Noha Radwan","Abhinav Valada","Wolfram Burgard"],"abstract":"Semantic understanding and localization are fundamental enablers of robot\nautonomy that have for the most part been tackled as disjoint problems. While\ndeep learning has enabled recent breakthroughs across a wide spectrum of scene\nunderstanding tasks, its applicability to state estimation tasks has been\nlimited due to the direct formulation that renders it incapable of encoding\nscene-specific constrains. In this work, we propose the VLocNet++ architecture\nthat employs a multitask learning approach to exploit the inter-task\nrelationship between learning semantics, regressing 6-DoF global pose and\nodometry, for the mutual benefit of each of these tasks. Our network overcomes\nthe aforementioned limitation by simultaneously embedding geometric and\nsemantic knowledge of the world into the pose regression network. We propose a\nnovel adaptive weighted fusion layer to aggregate motion-specific temporal\ninformation and to fuse semantic features into the localization stream based on\nregion activations. Furthermore, we propose a self-supervised warping technique\nthat uses the relative motion to warp intermediate network representations in\nthe segmentation stream for learning consistent semantics. Finally, we\nintroduce a first-of-a-kind urban outdoor localization dataset with pixel-level\nsemantic labels and multiple loops for training deep networks. Extensive\nexperiments on the challenging Microsoft 7-Scenes benchmark and our DeepLoc\ndataset demonstrate that our approach exceeds the state-of-the-art\noutperforming local feature-based methods while simultaneously performing\nmultiple tasks and exhibiting substantial robustness in challenging scenarios.","url_abs":"http://arxiv.org/abs/1804.08366v6","url_pdf":"http://arxiv.org/pdf/1804.08366v6.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"outdoor-localization","task_name":"Outdoor Localization"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"state-estimation","task_name":"State Estimation"},{"task_slug":"visual-localization","task_name":"Visual Localization"}],"methods":[],"datasets_introduced":[{"slug":"deeploc","name":"DeepLoc","full_name":"DeepLoc"},{"slug":"deeploccross","name":"DeepLocCross","full_name":"DeepLocCross"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.08366","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}