{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unim-ov3d-uni-modality-open-vocabulary-3d","title":"UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation","arxiv_id":"2401.11395","date":"2024-01-21","proceeding":null,"authors":["Qingdong He","Jinlong Peng","Zhengkai Jiang","Kai Wu","Xiaozhong Ji","Jiangning Zhang","Yabiao Wang","Chengjie Wang","Mingang Chen","Yunsheng Wu"],"abstract":"3D open-vocabulary scene understanding aims to recognize arbitrary novel categories beyond the base label space. However, existing works not only fail to fully utilize all the available modal information in the 3D domain but also lack sufficient granularity in representing the features of each modality. In this paper, we propose a unified multimodal 3D open-vocabulary scene understanding network, namely UniM-OV3D, which aligns point clouds with image, language and depth. To better integrate global and local features of the point clouds, we design a hierarchical point cloud feature extraction module that learns comprehensive fine-grained feature representations. Further, to facilitate the learning of coarse-to-fine point-semantic representations from captions, we propose the utilization of hierarchical 3D caption pairs, capitalizing on geometric constraints across various viewpoints of 3D scenes. Extensive experimental results demonstrate the effectiveness and superiority of our method in open-vocabulary semantic and instance segmentation, which achieves state-of-the-art performance on both indoor and outdoor benchmarks such as ScanNet, ScanNet200, S3IDS and nuScenes. Code is available at https://github.com/hithqd/UniM-OV3D.","url_abs":"https://arxiv.org/abs/2401.11395v3","url_pdf":"https://arxiv.org/pdf/2401.11395v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unim-ov3d-uni-modality-open-vocabulary-3d","repo_url":"https://github.com/hithqd/unim-ov3d","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[{"method_slug":"base","method_name":"BASE"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2401.11395","atlas_url":"https://app.syntology.ai/?focus=2401.11395","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2401.11395"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/hithqd/unim-ov3d","reach":null}],"summary":{"ran":1,"unverified":2},"by_repo_kind":{"official":{"samples":3,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"83afeeacb9f1ca25","entry":"SparseUNetTextSeg","repo":"hithqd/unim-ov3d","repo_kind":"official","path":"pcseg/models/vision_networks/sparseunet_textseg.py","file_url":"https://github.com/hithqd/unim-ov3d/blob/HEAD/pcseg/models/vision_networks/sparseunet_textseg.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"83afeeacb9f1ca25"}},{"code_sha256_prefix":"c56bbe1cc3df19fa","entry":"ModelTemplate","repo":"hithqd/unim-ov3d","repo_kind":"official","path":"pcseg/models/vision_networks/sparseunet_textseg.py","file_url":"https://github.com/hithqd/unim-ov3d/blob/HEAD/pcseg/models/vision_networks/sparseunet_textseg.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c56bbe1cc3df19fa"}},{"code_sha256_prefix":"4736609cee75702a","entry":"find_all_spconv_keys","repo":"hithqd/unim-ov3d","repo_kind":"official","path":"pcseg/models/vision_networks/sparseunet_textseg.py","file_url":"https://github.com/hithqd/unim-ov3d/blob/HEAD/pcseg/models/vision_networks/sparseunet_textseg.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4736609cee75702a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}