{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/oneformer3d-one-transformer-for-unified-point","title":"OneFormer3D: One Transformer for Unified Point Cloud Segmentation","arxiv_id":"2311.14405","date":"2023-11-24","proceeding":"CVPR 2024 1","authors":["Maxim Kolodiazhnyi","Anna Vorontsova","Anton Konushin","Danila Rukhovich"],"abstract":"Semantic, instance, and panoptic segmentation of 3D point clouds have been addressed using task-specific models of distinct design. Thereby, the similarity of all segmentation tasks and the implicit relationship between them have not been utilized effectively. This paper presents a unified, simple, and effective model addressing all these tasks jointly. The model, named OneFormer3D, performs instance and semantic segmentation consistently, using a group of learnable kernels, where each kernel is responsible for generating a mask for either an instance or a semantic category. These kernels are trained with a transformer-based decoder with unified instance and semantic queries passed as an input. Such a design enables training a model end-to-end in a single run, so that it achieves top performance on all three segmentation tasks simultaneously. Specifically, our OneFormer3D ranks 1st and sets a new state-of-the-art (+2.1 mAP50) in the ScanNet test leaderboard. We also demonstrate the state-of-the-art results in semantic, instance, and panoptic segmentation of ScanNet (+21 PQ), ScanNet200 (+3.8 mAP50), and S3DIS (+0.8 mIoU) datasets.","url_abs":"https://arxiv.org/abs/2311.14405v1","url_pdf":"https://arxiv.org/pdf/2311.14405v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"oneformer3d-one-transformer-for-unified-point","repo_url":"https://github.com/oneformer3d/oneformer3d","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"3d-instance-segmentation-1","task_name":"3D Instance Segmentation"},{"task_slug":"3d-object-detection","task_name":"3D Object Detection"},{"task_slug":"3d-semantic-segmentation","task_name":"3D Semantic Segmentation"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"panoptic-segmentation","task_name":"Panoptic Segmentation"},{"task_slug":"point-cloud-segmentation","task_name":"Point Cloud Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-instance-segmentation-on-s3dis","task":"3D Instance Segmentation","dataset":"S3DIS","model":"OneFormer3D","rank_in_archive_order":1,"of":21,"metrics":{"AP@50":"75.8","mAP":"63.0","mPrec":"82.3","mRec":"74.1"},"uses_additional_data":false},{"leaderboard":"/sota/3d-instance-segmentation-on-scannetv2","task":"3D Instance Segmentation","dataset":"ScanNet(v2)","model":"OneFromer3D","rank_in_archive_order":4,"of":32,"metrics":{"mAP":"56.6","mAP @ 50":"80.1","mAP@25":"89.6"},"uses_additional_data":false},{"leaderboard":"/sota/3d-object-detection-on-scannetv2","task":"3D Object Detection","dataset":"ScanNetV2","model":"OneFormer3D","rank_in_archive_order":5,"of":33,"metrics":{"mAP@0.25":"76.9","mAP@0.5":"65.3"},"uses_additional_data":false},{"leaderboard":"/sota/3d-semantic-segmentation-on-s3dis","task":"3D Semantic Segmentation","dataset":"S3DIS","model":"OneFormer3D","rank_in_archive_order":1,"of":6,"metrics":{"mIoU (6-Fold)":"75.0","mIoU (Area-5)":"72.4"},"uses_additional_data":false},{"leaderboard":"/sota/3d-semantic-segmentation-on-scannet200","task":"3D Semantic Segmentation","dataset":"ScanNet200","model":"OneFormer3D","rank_in_archive_order":13,"of":16,"metrics":{"val mIoU":"30.1"},"uses_additional_data":false},{"leaderboard":"/sota/panoptic-segmentation-on-scannet","task":"Panoptic Segmentation","dataset":"ScanNet","model":"OneFormer3D","rank_in_archive_order":1,"of":4,"metrics":{"PQ":"71.2","PQ_st":"86.1","PQ_th":"69.6"},"uses_additional_data":false},{"leaderboard":"/sota/panoptic-segmentation-on-scannetv2","task":"Panoptic Segmentation","dataset":"ScanNetV2","model":"OneFormer3D","rank_in_archive_order":1,"of":5,"metrics":{"PQ":"71.2"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-segmentation-on-scannet","task":"Semantic Segmentation","dataset":"ScanNet","model":"OneFormer3D","rank_in_archive_order":11,"of":45,"metrics":{"val mIoU":"76.6"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2311.14405","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}