{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/imvotenet-boosting-3d-object-detection-in","title":"ImVoteNet: Boosting 3D Object Detection in Point Clouds with Image Votes","arxiv_id":"2001.10692","date":"2020-01-29","proceeding":"CVPR 2020 6","authors":["Charles R. Qi","Xinlei Chen","Or Litany","Leonidas J. Guibas"],"abstract":"3D object detection has seen quick progress thanks to advances in deep learning on point clouds. A few recent works have even shown state-of-the-art performance with just point clouds input (e.g. VoteNet). However, point cloud data have inherent limitations. They are sparse, lack color information and often suffer from sensor noise. Images, on the other hand, have high resolution and rich texture. Thus they can complement the 3D geometry provided by point clouds. Yet how to effectively use image information to assist point cloud based detection is still an open question. In this work, we build on top of VoteNet and propose a 3D detection architecture called ImVoteNet specialized for RGB-D scenes. ImVoteNet is based on fusing 2D votes in images and 3D votes in point clouds. Compared to prior work on multi-modal detection, we explicitly extract both geometric and semantic features from the 2D images. We leverage camera parameters to lift these features to 3D. To improve the synergy of 2D-3D feature fusion, we also propose a multi-tower training scheme. We validate our model on the challenging SUN RGB-D dataset, advancing state-of-the-art results by 5.7 mAP. We also provide rich ablation studies to analyze the contribution of each design choice.","url_abs":"https://arxiv.org/abs/2001.10692v1","url_pdf":"https://arxiv.org/pdf/2001.10692v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"imvotenet-boosting-3d-object-detection-in","repo_url":"https://github.com/facebookresearch/imvotenet","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"3d-object-detection","task_name":"3D Object Detection"},{"task_slug":"3d-geometry","task_name":"3D geometry"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"open-question","task_name":"Open-Ended Question Answering"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-object-detection-on-sun-rgbd","task":"3D Object Detection","dataset":"SUN-RGBD","model":"ImVoteNet","rank_in_archive_order":3,"of":8,"metrics":{"mAP@0.25":"63.4"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2001.10692","atlas_url":"https://app.syntology.ai/?focus=2001.10692","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2001.10692"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/facebookresearch/imvotenet","reach":null}],"summary":{"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"847f37b80100244c","entry":"sample_valid_seeds","repo":"facebookresearch/imvotenet","repo_kind":"official","path":"models/imvotenet.py","file_url":"https://github.com/facebookresearch/imvotenet/blob/HEAD/models/imvotenet.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"847f37b80100244c"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}