{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/language-grounded-indoor-3d-semantic","title":"Language-Grounded Indoor 3D Semantic Segmentation in the Wild","arxiv_id":"2204.07761","date":"2022-04-16","proceeding":null,"authors":["David Rozenberszki","Or Litany","Angela Dai"],"abstract":"Recent advances in 3D semantic segmentation with deep neural networks have shown remarkable success, with rapid performance increase on available datasets. However, current 3D semantic segmentation benchmarks contain only a small number of categories -- less than 30 for ScanNet and SemanticKITTI, for instance, which are not enough to reflect the diversity of real environments (e.g., semantic image understanding covers hundreds to thousands of classes). Thus, we propose to study a larger vocabulary for 3D semantic segmentation with a new extended benchmark on ScanNet data with 200 class categories, an order of magnitude more than previously studied. This large number of class categories also induces a large natural class imbalance, both of which are challenging for existing 3D semantic segmentation methods. To learn more robust 3D features in this context, we propose a language-driven pre-training method to encourage learned 3D features that might have limited training examples to lie close to their pre-trained text embeddings. Extensive experiments show that our approach consistently outperforms state-of-the-art 3D pre-training for 3D semantic segmentation on our proposed benchmark (+9% relative mIoU), including limited-data scenarios with +25% relative mIoU using only 5% annotations.","url_abs":"https://arxiv.org/abs/2204.07761v2","url_pdf":"https://arxiv.org/pdf/2204.07761v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"language-grounded-indoor-3d-semantic","repo_url":"https://github.com/RozDavid/LanguageGroundedSemseg","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"3d-semantic-segmentation","task_name":"3D Semantic Segmentation"},{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[{"slug":"scannet200","name":"ScanNet200","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-semantic-segmentation-on-scannet200","task":"3D Semantic Segmentation","dataset":"ScanNet200","model":"LGround","rank_in_archive_order":14,"of":16,"metrics":{"test mIoU":"27.2","val mIoU":"28.8"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2204.07761","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2204.07761"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/RozDavid/LanguageGroundedSemseg","reach":null}],"summary":{"ran":2},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"54dc2f49b405794b","entry":"AttributeFittingModel","repo":"RozDavid/LanguageGroundedSemseg","repo_kind":"official","path":"lib/losses/ContrastiveLanguageLoss.py","file_url":"https://github.com/RozDavid/LanguageGroundedSemseg/blob/HEAD/lib/losses/ContrastiveLanguageLoss.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"54dc2f49b405794b"}},{"code_sha256_prefix":"28ad76ee510ae152","entry":"ContrastiveLanguageLoss","repo":"RozDavid/LanguageGroundedSemseg","repo_kind":"official","path":"lib/losses/ContrastiveLanguageLoss.py","file_url":"https://github.com/RozDavid/LanguageGroundedSemseg/blob/HEAD/lib/losses/ContrastiveLanguageLoss.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"28ad76ee510ae152"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}