{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dels-3d-deep-localization-and-segmentation","title":"DeLS-3D: Deep Localization and Segmentation with a 3D Semantic Map","arxiv_id":"1805.04949","date":"2018-05-13","proceeding":"CVPR 2018 6","authors":["Peng Wang","Ruigang Yang","Binbin Cao","Wei Xu","Yuanqing Lin"],"abstract":"For applications such as autonomous driving, self-localization/camera pose\nestimation and scene parsing are crucial technologies. In this paper, we\npropose a unified framework to tackle these two problems simultaneously. The\nuniqueness of our design is a sensor fusion scheme which integrates camera\nvideos, motion sensors (GPS/IMU), and a 3D semantic map in order to achieve\nrobustness and efficiency of the system. Specifically, we first have an initial\ncoarse camera pose obtained from consumer-grade GPS/IMU, based on which a label\nmap can be rendered from the 3D semantic map. Then, the rendered label map and\nthe RGB image are jointly fed into a pose CNN, yielding a corrected camera\npose. In addition, to incorporate temporal information, a multi-layer recurrent\nneural network (RNN) is further deployed improve the pose accuracy. Finally,\nbased on the pose from RNN, we render a new label map, which is fed together\nwith the RGB image into a segment CNN which produces per-pixel semantic label.\nIn order to validate our approach, we build a dataset with registered 3D point\nclouds and video camera images. Both the point clouds and the images are\nsemantically-labeled. Each video frame has ground truth pose from highly\naccurate motion sensors. We show that practically, pose estimation solely\nrelying on images like PoseNet may fail due to street view confusion, and it is\nimportant to fuse multiple sensors. Finally, various ablation studies are\nperformed, which demonstrate the effectiveness of the proposed system. In\nparticular, we show that scene parsing and pose estimation are mutually\nbeneficial to achieve a more robust and accurate system.","url_abs":"http://arxiv.org/abs/1805.04949v1","url_pdf":"http://arxiv.org/pdf/1805.04949v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dels-3d-deep-localization-and-segmentation","repo_url":"https://github.com/pengwangucla/DeLS-3D","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"camera-pose-estimation","task_name":"Camera Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"scene-parsing","task_name":"Scene Parsing"},{"task_slug":"sensor-fusion","task_name":"Sensor Fusion"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.04949","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}