{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mvx-net-multimodal-voxelnet-for-3d-object","title":"MVX-Net: Multimodal VoxelNet for 3D Object Detection","arxiv_id":"1904.01649","date":"2019-04-02","proceeding":null,"authors":["Vishwanath A. Sindagi","Yin Zhou","Oncel Tuzel"],"abstract":"Many recent works on 3D object detection have focused on designing neural\nnetwork architectures that can consume point cloud data. While these approaches\ndemonstrate encouraging performance, they are typically based on a single\nmodality and are unable to leverage information from other modalities, such as\na camera. Although a few approaches fuse data from different modalities, these\nmethods either use a complicated pipeline to process the modalities\nsequentially, or perform late-fusion and are unable to learn interaction\nbetween different modalities at early stages. In this work, we present\nPointFusion and VoxelFusion: two simple yet effective early-fusion approaches\nto combine the RGB and point cloud modalities, by leveraging the recently\nintroduced VoxelNet architecture. Evaluation on the KITTI dataset demonstrates\nsignificant improvements in performance over approaches which only use point\ncloud data. Furthermore, the proposed method provides results competitive with\nthe state-of-the-art multimodal algorithms, achieving top-2 ranking in five of\nthe six bird's eye view and 3D detection categories on the KITTI benchmark, by\nusing a simple single stage network.","url_abs":"http://arxiv.org/abs/1904.01649v1","url_pdf":"http://arxiv.org/pdf/1904.01649v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mvx-net-multimodal-voxelnet-for-3d-object","repo_url":"https://github.com/open-mmlab/mmdetection3d","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"3d-object-detection","task_name":"3D Object Detection"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-object-detection-on-dair-v2x-i","task":"3D Object Detection","dataset":"DAIR-V2X-I","model":"MVXNet","rank_in_archive_order":7,"of":9,"metrics":{"AP|R40(easy)":"71.0","AP|R40(hard)":"53.8","AP|R40(moderate)":"53.7"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.01649","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}