{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pointclip-point-cloud-understanding-by-clip","title":"PointCLIP: Point Cloud Understanding by CLIP","arxiv_id":"2112.02413","date":"2021-12-04","proceeding":"CVPR 2022 1","authors":["Renrui Zhang","Ziyu Guo","Wei zhang","Kunchang Li","Xupeng Miao","Bin Cui","Yu Qiao","Peng Gao","Hongsheng Li"],"abstract":"Recently, zero-shot and few-shot learning via Contrastive Vision-Language Pre-training (CLIP) have shown inspirational performance on 2D visual recognition, which learns to match images with their corresponding texts in open-vocabulary settings. However, it remains under explored that whether CLIP, pre-trained by large-scale image-text pairs in 2D, can be generalized to 3D recognition. In this paper, we identify such a setting is feasible by proposing PointCLIP, which conducts alignment between CLIP-encoded point cloud and 3D category texts. Specifically, we encode a point cloud by projecting it into multi-view depth maps without rendering, and aggregate the view-wise zero-shot prediction to achieve knowledge transfer from 2D to 3D. On top of that, we design an inter-view adapter to better extract the global feature and adaptively fuse the few-shot knowledge learned from 3D into CLIP pre-trained in 2D. By just fine-tuning the lightweight adapter in the few-shot settings, the performance of PointCLIP could be largely improved. In addition, we observe the complementary property between PointCLIP and classical 3D-supervised networks. By simple ensembling, PointCLIP boosts baseline's performance and even surpasses state-of-the-art models. Therefore, PointCLIP is a promising alternative for effective 3D point cloud understanding via CLIP under low resource cost and data regime. We conduct thorough experiments on widely-adopted ModelNet10, ModelNet40 and the challenging ScanObjectNN to demonstrate the effectiveness of PointCLIP. The code is released at https://github.com/ZrrSkywalker/PointCLIP.","url_abs":"https://arxiv.org/abs/2112.02413v1","url_pdf":"https://arxiv.org/pdf/2112.02413v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"pointclip-point-cloud-understanding-by-clip","repo_url":"https://github.com/zrrskywalker/pointclip","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"pointclip-point-cloud-understanding-by-clip","repo_url":"https://github.com/pku-dair/hetu","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"3d-open-vocabulary-instance-segmentation","task_name":"3D Open-Vocabulary Instance Segmentation"},{"task_slug":"few-shot-learning","task_name":"Few-Shot Learning"},{"task_slug":"open-vocabulary-object-detection","task_name":"Open Vocabulary Object Detection"},{"task_slug":"training-free-3d-part-segmentation","task_name":"Training-free 3D Part Segmentation"},{"task_slug":"training-free-3d-point-cloud-classification","task_name":"Training-free 3D Point Cloud Classification"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"zero-shot-transfer-3d-point-cloud","task_name":"Zero-Shot Transfer 3D Point Cloud Classification"},{"task_slug":"zero-shot-3d-point-cloud-classification","task_name":"Zero-shot 3D Point Cloud Classification"},{"task_slug":"zero-shot-3d-classification","task_name":"Zero-shot 3D classification"}],"methods":[{"method_slug":"adapter","method_name":"Adapter"},{"method_slug":"clip","method_name":"CLIP"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-open-vocabulary-instance-segmentation-on-3","task":"3D Open-Vocabulary Instance Segmentation","dataset":"STPLS3D","model":"PointCLIP","rank_in_archive_order":3,"of":3,"metrics":{"AP50":"02.6"},"uses_additional_data":false},{"leaderboard":"/sota/training-free-3d-part-segmentation-on","task":"Training-free 3D Part Segmentation","dataset":"ShapeNet-Part","model":"PointCLIP","rank_in_archive_order":3,"of":3,"metrics":{"Need 3D Data?":"No","mIoU":"31.0"},"uses_additional_data":true},{"leaderboard":"/sota/training-free-3d-point-cloud-classification","task":"Training-free 3D Point Cloud Classification","dataset":"ModelNet40","model":"PointCLIP","rank_in_archive_order":7,"of":7,"metrics":{"Accuracy (%)":"20.2","Need 3D Data?":"No"},"uses_additional_data":true},{"leaderboard":"/sota/training-free-3d-point-cloud-classification-1","task":"Training-free 3D Point Cloud Classification","dataset":"ScanObjectNN","model":"PointCLIP","rank_in_archive_order":6,"of":6,"metrics":{"Accuracy (%)":"15.4","Need 3D Data?":"No"},"uses_additional_data":true},{"leaderboard":"/sota/zero-shot-transfer-3d-point-cloud-1","task":"Zero-Shot Transfer 3D Point Cloud Classification","dataset":"ModelNet10","model":"PointCLIP","rank_in_archive_order":4,"of":4,"metrics":{"Accuracy (%)":"30.23"},"uses_additional_data":true},{"leaderboard":"/sota/zero-shot-transfer-3d-point-cloud","task":"Zero-Shot Transfer 3D Point Cloud Classification","dataset":"ModelNet40","model":"PointCLIP","rank_in_archive_order":16,"of":16,"metrics":{"Accuracy (%)":"20.18"},"uses_additional_data":true},{"leaderboard":"/sota/zero-shot-transfer-3d-point-cloud-2","task":"Zero-Shot Transfer 3D Point Cloud Classification","dataset":"ScanObjectNN","model":"PointCLIP","rank_in_archive_order":10,"of":10,"metrics":{"OBJ_BG Accuracy(%)":"21.34","OBJ_ONLY Accuracy(%)":"19.28","PB_T50_RS Accuracy (%)":"15.38"},"uses_additional_data":true}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2112.02413","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}