{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-3d-representations-from-2d-pre","title":"Learning 3D Representations from 2D Pre-trained Models via Image-to-Point Masked Autoencoders","arxiv_id":"2212.06785","date":"2022-12-13","proceeding":"CVPR 2023 1","authors":["Renrui Zhang","Liuhui Wang","Yu Qiao","Peng Gao","Hongsheng Li"],"abstract":"Pre-training by numerous image data has become de-facto for robust 2D representations. In contrast, due to the expensive data acquisition and annotation, a paucity of large-scale 3D datasets severely hinders the learning for high-quality 3D features. In this paper, we propose an alternative to obtain superior 3D representations from 2D pre-trained models via Image-to-Point Masked Autoencoders, named as I2P-MAE. By self-supervised pre-training, we leverage the well learned 2D knowledge to guide 3D masked autoencoding, which reconstructs the masked point tokens with an encoder-decoder architecture. Specifically, we first utilize off-the-shelf 2D models to extract the multi-view visual features of the input point cloud, and then conduct two types of image-to-point learning schemes on top. For one, we introduce a 2D-guided masking strategy that maintains semantically important point tokens to be visible for the encoder. Compared to random masking, the network can better concentrate on significant 3D structures and recover the masked tokens from key spatial cues. For another, we enforce these visible tokens to reconstruct the corresponding multi-view 2D features after the decoder. This enables the network to effectively inherit high-level 2D semantics learned from rich image data for discriminative 3D modeling. Aided by our image-to-point pre-training, the frozen I2P-MAE, without any fine-tuning, achieves 93.4% accuracy for linear SVM on ModelNet40, competitive to the fully trained results of existing methods. By further fine-tuning on on ScanObjectNN's hardest split, I2P-MAE attains the state-of-the-art 90.11% accuracy, +3.68% to the second-best, demonstrating superior transferable capacity. Code will be available at https://github.com/ZrrSkywalker/I2P-MAE.","url_abs":"https://arxiv.org/abs/2212.06785v1","url_pdf":"https://arxiv.org/pdf/2212.06785v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-3d-representations-from-2d-pre","repo_url":"https://github.com/zrrskywalker/i2p-mae","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"learning-3d-representations-from-2d-pre","repo_url":"https://github.com/zrrskywalker/point-m2ae","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"3d-point-cloud-classification","task_name":"3D Point Cloud Classification"},{"task_slug":"3d-point-cloud-linear-classification","task_name":"3D Point Cloud Linear Classification"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"few-shot-3d-point-cloud-classification","task_name":"Few-Shot 3D Point Cloud Classification"}],"methods":[{"method_slug":"svm","method_name":"SVM"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-point-cloud-classification-on-scanobjectnn","task":"3D Point Cloud Classification","dataset":"ScanObjectNN","model":"I2P-MAE (no voting)","rank_in_archive_order":21,"of":77,"metrics":{"OBJ-BG (OA)":"94.15","OBJ-ONLY (OA)":"91.57","Overall Accuracy":"90.11"},"uses_additional_data":true},{"leaderboard":"/sota/3d-point-cloud-linear-classification-on","task":"3D Point Cloud Linear Classification","dataset":"ModelNet40","model":"I2P-MAE","rank_in_archive_order":3,"of":20,"metrics":{"Overall Accuracy":"93.4"},"uses_additional_data":true},{"leaderboard":"/sota/few-shot-3d-point-cloud-classification-on-3","task":"Few-Shot 3D Point Cloud Classification","dataset":"ModelNet40 10-way (10-shot)","model":"I2P-MAE","rank_in_archive_order":15,"of":31,"metrics":{"Overall Accuracy":"92.6","Standard Deviation":"5.0"},"uses_additional_data":true},{"leaderboard":"/sota/few-shot-3d-point-cloud-classification-on-4","task":"Few-Shot 3D Point Cloud Classification","dataset":"ModelNet40 10-way (20-shot)","model":"I2P-MAE","rank_in_archive_order":12,"of":31,"metrics":{"Overall Accuracy":"95.5","Standard Deviation":"3.0"},"uses_additional_data":true},{"leaderboard":"/sota/few-shot-3d-point-cloud-classification-on-1","task":"Few-Shot 3D Point Cloud Classification","dataset":"ModelNet40 5-way (10-shot)","model":"I2P-MAE","rank_in_archive_order":11,"of":30,"metrics":{"Overall Accuracy":"97.0","Standard Deviation":"1.8"},"uses_additional_data":true},{"leaderboard":"/sota/few-shot-3d-point-cloud-classification-on-2","task":"Few-Shot 3D Point Cloud Classification","dataset":"ModelNet40 5-way (20-shot)","model":"I2P-MAE","rank_in_archive_order":10,"of":30,"metrics":{"Overall Accuracy":"98.3","Standard Deviation":"1.3"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2212.06785","atlas_url":"https://app.syntology.ai/?focus=2212.06785","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2212.06785"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zrrskywalker/point-m2ae","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zrrskywalker/i2p-mae","reach":{"status":"ok"}}],"summary":{"ran_fixture":2,"ran":2,"ran_honours":1,"unverified":1},"by_repo_kind":{"listed":{"samples":6,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"449a0265144f6530","entry":"index_points","repo":"zrrskywalker/point-m2ae","repo_kind":"listed","path":"models/modules.py","file_url":"https://github.com/zrrskywalker/point-m2ae/blob/HEAD/models/modules.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":2,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"449a0265144f6530"}},{"code_sha256_prefix":"f80066a00e7156a2","entry":"farthest_point_sample","repo":"zrrskywalker/point-m2ae","repo_kind":"listed","path":"datasets/ModelNetDataset.py","file_url":"https://github.com/zrrskywalker/point-m2ae/blob/HEAD/datasets/ModelNetDataset.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f80066a00e7156a2"}},{"code_sha256_prefix":"e609e218822ed309","entry":"load_modelnet_data","repo":"zrrskywalker/point-m2ae","repo_kind":"listed","path":"datasets/ModelNetDataset.py","file_url":"https://github.com/zrrskywalker/point-m2ae/blob/HEAD/datasets/ModelNetDataset.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e609e218822ed309"}},{"code_sha256_prefix":"4783fbece52f500e","entry":"pc_normalize","repo":"zrrskywalker/point-m2ae","repo_kind":"listed","path":"datasets/ModelNetDataset.py","file_url":"https://github.com/zrrskywalker/point-m2ae/blob/HEAD/datasets/ModelNetDataset.py","link_basis":"harvester_set","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4783fbece52f500e"}},{"code_sha256_prefix":"3bfe172e686075cd","entry":"square_distance","repo":"zrrskywalker/point-m2ae","repo_kind":"listed","path":"models/modules.py","file_url":"https://github.com/zrrskywalker/point-m2ae/blob/HEAD/models/modules.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3bfe172e686075cd"}},{"code_sha256_prefix":"b1227ddb721e2999","entry":"timeit","repo":"zrrskywalker/point-m2ae","repo_kind":"listed","path":"segmentation/pointnet2_utils.py","file_url":"https://github.com/zrrskywalker/point-m2ae/blob/HEAD/segmentation/pointnet2_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b1227ddb721e2999"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}