{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/monocular-semantic-occupancy-grid-mapping","title":"Monocular Semantic Occupancy Grid Mapping with Convolutional Variational Encoder-Decoder Networks","arxiv_id":"1804.02176","date":"2018-04-06","proceeding":null,"authors":["Chenyang Lu","Marinus Jacobus Gerardus van de Molengraft","Gijs Dubbelman"],"abstract":"In this work, we research and evaluate end-to-end learning of monocular\nsemantic-metric occupancy grid mapping from weak binocular ground truth. The\nnetwork learns to predict four classes, as well as a camera to bird's eye view\nmapping. At the core, it utilizes a variational encoder-decoder network that\nencodes the front-view visual information of the driving scene and subsequently\ndecodes it into a 2-D top-view Cartesian coordinate system. The evaluations on\nCityscapes show that the end-to-end learning of semantic-metric occupancy grids\noutperforms the deterministic mapping approach with flat-plane assumption by\nmore than 12% mean IoU. Furthermore, we show that the variational sampling with\na relatively small embedding vector brings robustness against vehicle dynamic\nperturbations, and generalizability for unseen KITTI data. Our network achieves\nreal-time inference rates of approx. 35 Hz for an input image with a resolution\nof 256x512 pixels and an output map with 64x64 occupancy grid cells using a\nTitan V GPU.","url_abs":"http://arxiv.org/abs/1804.02176v3","url_pdf":"http://arxiv.org/pdf/1804.02176v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"bird-s-eye-view-semantic-segmentation","task_name":"Bird's-Eye View Semantic Segmentation"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":null,"task_name":"GPU"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/bird-s-eye-view-semantic-segmentation-on","task":"Bird's-Eye View Semantic Segmentation","dataset":"nuScenes","model":"VED","rank_in_archive_order":17,"of":17,"metrics":{"IoU veh - 224x480 - No vis filter - 100x50 at 0.25":"8.8"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.02176","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}