{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/geometry-aware-learning-of-maps-for-camera","title":"Geometry-Aware Learning of Maps for Camera Localization","arxiv_id":"1712.03342","date":"2017-12-09","proceeding":"CVPR 2018 6","authors":["Samarth Brahmbhatt","Jinwei Gu","Kihwan Kim","James Hays","Jan Kautz"],"abstract":"Maps are a key component in image-based camera localization and visual SLAM\nsystems: they are used to establish geometric constraints between images,\ncorrect drift in relative pose estimation, and relocalize cameras after lost\ntracking. The exact definitions of maps, however, are often\napplication-specific and hand-crafted for different scenarios (e.g. 3D\nlandmarks, lines, planes, bags of visual words). We propose to represent maps\nas a deep neural net called MapNet, which enables learning a data-driven map\nrepresentation. Unlike prior work on learning maps, MapNet exploits cheap and\nubiquitous sensory inputs like visual odometry and GPS in addition to images\nand fuses them together for camera localization. Geometric constraints\nexpressed by these inputs, which have traditionally been used in bundle\nadjustment or pose-graph optimization, are formulated as loss terms in MapNet\ntraining and also used during inference. In addition to directly improving\nlocalization accuracy, this allows us to update the MapNet (i.e., maps) in a\nself-supervised manner using additional unlabeled video sequences from the\nscene. We also propose a novel parameterization for camera rotation which is\nbetter suited for deep-learning based camera pose regression. Experimental\nresults on both the indoor 7-Scenes dataset and the outdoor Oxford RobotCar\ndataset show significant performance improvement over prior work. The MapNet\nproject webpage is https://goo.gl/mRB3Au.","url_abs":"http://arxiv.org/abs/1712.03342v3","url_pdf":"http://arxiv.org/pdf/1712.03342v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"geometry-aware-learning-of-maps-for-camera","repo_url":"https://github.com/NVlabs/geomapnet","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"camera-localization","task_name":"Camera Localization"},{"task_slug":"visual-localization","task_name":"Visual Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/camera-localization-on-oxford-robotcar-full","task":"Camera Localization","dataset":"Oxford RobotCar Full","model":"MapNet++","rank_in_archive_order":5,"of":6,"metrics":{"Mean Translation Error":"29.5"},"uses_additional_data":false},{"leaderboard":"/sota/visual-localization-on-oxford-radar-robotcar","task":"Visual Localization","dataset":"Oxford Radar RobotCar (Full-6)","model":"MapNet","rank_in_archive_order":16,"of":16,"metrics":{"Mean Translation Error":"48.21"},"uses_additional_data":false},{"leaderboard":"/sota/visual-localization-on-oxford-robotcar-full","task":"Visual Localization","dataset":"Oxford RobotCar Full","model":"MapNet++","rank_in_archive_order":5,"of":6,"metrics":{"Mean Translation Error":"29.5"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1712.03342","atlas_url":"https://app.syntology.ai/?focus=1712.03342","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}