{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rethinking-masked-representation-learning-for","title":"Rethinking Masked Representation Learning for 3D Point Cloud Understanding","arxiv_id":null,"date":"2024-12-26","proceeding":"IEEE Transactions on Image Processing 2024 12","authors":["Chuxin Wang","Yixin Zha","Jianfeng He","Wenfei Yang","Tianzhu Zhang"],"abstract":"Self-supervised point cloud representation learning aims to acquire robust and general feature representations from unlabeled data. Recently, masked point modeling-based methods have shown significant performance improvements for point cloud understanding, yet these methods rely on overlapping grouping strategies (k-nearest neighbor algorithm) resulting in early leakage of structural information of mask groups, and overlook the semantic modeling of object components resulting in parts with the same semantics having obvious feature differences due to position differences. In this work, we rethink grouping strategies and pretext tasks that are more suitable for self-supervised point cloud representation learning and propose a novel hierarchical masked representation learning method, including an optimal transport-based hierarchical grouping strategy, a prototype-based part modeling module, and a hierarchical attention encoder. The proposed method enjoys several merits. First, the proposed grouping strategy partitions the point cloud into non-overlapping groups, eliminating the early leakage of structural information in the masked groups. Second, the proposed prototype-based part modeling module dynamically models different object components, ensuring feature consistency on parts with the same semantics. Extensive experiments on four downstream tasks demonstrate that our method surpasses state-of-the-art 3D representation learning methods. Comprehensive ablation studies and visualizations demonstrate the effectiveness of the proposed modules.","url_abs":"https://ieeexplore.ieee.org/document/10815033","url_pdf":"https://ieeexplore.ieee.org/document/10815033","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rethinking-masked-representation-learning-for","repo_url":"https://github.com/OpenSpaceAI/OTMae3D","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"3d-part-segmentation","task_name":"3D Part Segmentation"},{"task_slug":"3d-point-cloud-classification","task_name":"3D Point Cloud Classification"},{"task_slug":"few-shot-3d-point-cloud-classification","task_name":"Few-Shot 3D Point Cloud Classification"},{"task_slug":"representation-learning","task_name":"Representation Learning"}],"methods":[{"method_slug":"attention","method_name":"Attention"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-part-segmentation-on-shapenet-part","task":"3D Part Segmentation","dataset":"ShapeNet-Part","model":"OTMae3D","rank_in_archive_order":12,"of":67,"metrics":{"Class Average IoU":"85.1","Instance Average IoU":"86.8"},"uses_additional_data":true},{"leaderboard":"/sota/3d-point-cloud-classification-on-modelnet40","task":"3D Point Cloud Classification","dataset":"ModelNet40","model":"OTMae3D","rank_in_archive_order":13,"of":111,"metrics":{"Overall Accuracy":"94.5"},"uses_additional_data":true},{"leaderboard":"/sota/3d-point-cloud-classification-on-modelnet40","task":"3D Point Cloud Classification","dataset":"ModelNet40","model":"OTMae3D (w/o Voting)","rank_in_archive_order":16,"of":111,"metrics":{"Overall Accuracy":"94.3"},"uses_additional_data":true},{"leaderboard":"/sota/3d-point-cloud-classification-on-scanobjectnn","task":"3D Point Cloud Classification","dataset":"ScanObjectNN","model":"OTMae3D","rank_in_archive_order":31,"of":77,"metrics":{"FLOPs":"6.29","OBJ-BG (OA)":"92.9","OBJ-ONLY (OA)":"92.3","Overall Accuracy":"89.0"},"uses_additional_data":true},{"leaderboard":"/sota/few-shot-3d-point-cloud-classification-on-3","task":"Few-Shot 3D Point Cloud Classification","dataset":"ModelNet40 10-way (10-shot)","model":"OTMae3D","rank_in_archive_order":11,"of":31,"metrics":{"Overall Accuracy":"93.2","Standard Deviation":"3.4"},"uses_additional_data":true},{"leaderboard":"/sota/few-shot-3d-point-cloud-classification-on-4","task":"Few-Shot 3D Point Cloud Classification","dataset":"ModelNet40 10-way (20-shot)","model":"OTMae3D","rank_in_archive_order":10,"of":31,"metrics":{"Overall Accuracy":"95.6","Standard Deviation":"2.6"},"uses_additional_data":true},{"leaderboard":"/sota/few-shot-3d-point-cloud-classification-on-1","task":"Few-Shot 3D Point Cloud Classification","dataset":"ModelNet40 5-way (10-shot)","model":"OTMae3D","rank_in_archive_order":10,"of":30,"metrics":{"Overall Accuracy":"97.2","Standard Deviation":"2.3"},"uses_additional_data":true},{"leaderboard":"/sota/few-shot-3d-point-cloud-classification-on-2","task":"Few-Shot 3D Point Cloud Classification","dataset":"ModelNet40 5-way (20-shot)","model":"OTMae3D","rank_in_archive_order":8,"of":30,"metrics":{"Overall Accuracy":"98.7","Standard Deviation":"1.2"},"uses_additional_data":true}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}