Papers › PointLLM: Empowering Large Language Models to Understand Point Clouds

PointLLM: Empowering Large Language Models to Understand Point Clouds

31 Aug 2023arXiv:2308.16911archive 2025-07-28

Runsen Xu, Xiaolong Wang, Tai Wang, Yilun Chen, Jiangmiao Pang, Dahua Lin

The unprecedented advancements in Large Language Models (LLMs) have shown a profound impact on natural language processing but are yet to fully embrace the realm of 3D understanding. This paper introduces PointLLM, a preliminary effort to fill this gap, enabling LLMs to understand point clouds and offering a new avenue beyond 2D visual data. PointLLM understands colored object point clouds with human instructions and generates contextually appropriate responses, illustrating its grasp of point clouds and common sense. Specifically, it leverages a point cloud encoder with a powerful LLM to effectively fuse geometric, appearance, and linguistic information. We collect a novel dataset comprising 660K simple and 70K complex point-text instruction pairs to enable a two-stage training strategy: aligning latent spaces and subsequently instruction-tuning the unified model. To rigorously evaluate the perceptual and generalization capabilities of PointLLM, we establish two benchmarks: Generative 3D Object Classification and 3D Object Captioning, assessed through three different methods, including human evaluation, GPT-4/ChatGPT evaluation, and traditional metrics. Experimental results reveal PointLLM's superior performance over existing 2D and 3D baselines, with a notable achievement in human-evaluated object captioning tasks where it surpasses human annotators in over 50% of the samples. Codes, datasets, and benchmarks are available at https://github.com/OpenRobotLab/PointLLM .

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

openrobotlab/pointllm officialmentioned in papermentioned on GitHubpytorch report
Pointcept/GPT4Point mentioned on GitHubpytorchMIT report
qizekun/ShapeLLM mentioned on GitHubpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Object Captioning3D Object Classification3D Question Answering (3D-QA)Common Sense ReasoningGenerative 3D Object ClassificationObject

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D Object Captioning Objaverse PointLLM-13B V1.2 Sentence-BERT 47.91 #3 of 6 Archive leaderboard report
3D Object Captioning Objaverse PointLLM-13B V1.2 Correctness 3.10 #3 of 6 Archive leaderboard report
3D Object Captioning Objaverse PointLLM-13B V1.2 GPT-4 48.15 #3 of 6 Archive leaderboard report
3D Object Captioning Objaverse PointLLM-13B V1.2 Hallucination 0.84 #3 of 6 Archive leaderboard report
3D Object Captioning Objaverse PointLLM-13B V1.2 Precision 78.75 #3 of 6 Archive leaderboard report
3D Object Captioning Objaverse PointLLM-13B V1.2 SimCSE 49.12 #3 of 6 Archive leaderboard report
3D Object Captioning Objaverse PointLLM-7B V1.2 Sentence-BERT 47.47 #5 of 6 Archive leaderboard report
3D Object Captioning Objaverse PointLLM-7B V1.2 Correctness 3.04 #5 of 6 Archive leaderboard report
3D Object Captioning Objaverse PointLLM-7B V1.2 GPT-4 44.85 #5 of 6 Archive leaderboard report
3D Object Captioning Objaverse PointLLM-7B V1.2 Hallucination 0.66 #5 of 6 Archive leaderboard report
3D Object Captioning Objaverse PointLLM-7B V1.2 Precision 82.14 #5 of 6 Archive leaderboard report
3D Object Captioning Objaverse PointLLM-7B V1.2 SimCSE 48.55 #5 of 6 Archive leaderboard report
3D Question Answering (3D-QA) 3D MM-Vet PointLLM-13B v1.2 Overall Accuracy 46.6 #3 of 5 Archive leaderboard report
3D Question Answering (3D-QA) 3D MM-Vet PointLLM-7B v1.2 Overall Accuracy 41.2 #4 of 5 Archive leaderboard report
Generative 3D Object Classification ModelNet40 PointLLM-13B v1.2 ModelNet40 (Average) 52.78 #4 of 6 Archive leaderboard report
Generative 3D Object Classification ModelNet40 PointLLM-13B v1.2 ModelNet40 (C) 52.55 #4 of 6 Archive leaderboard report
Generative 3D Object Classification ModelNet40 PointLLM-13B v1.2 ModelNet40 (I) 53.00 #4 of 6 Archive leaderboard report
Generative 3D Object Classification ModelNet40 PointLLM-7B v1.2 ModelNet40 (Average) 52.63 #5 of 6 Archive leaderboard report
Generative 3D Object Classification ModelNet40 PointLLM-7B v1.2 ModelNet40 (C) 51.82 #5 of 6 Archive leaderboard report
Generative 3D Object Classification ModelNet40 PointLLM-7B v1.2 ModelNet40 (I) 53.44 #5 of 6 Archive leaderboard report
Generative 3D Object Classification Objaverse PointLLM-13B v1.2 Objaverse (Average) 54.00 #3 of 7 Archive leaderboard report
Generative 3D Object Classification Objaverse PointLLM-13B v1.2 Objaverse (C) 51.50 #3 of 7 Archive leaderboard report
Generative 3D Object Classification Objaverse PointLLM-13B v1.2 Objaverse (I) 56.50 #3 of 7 Archive leaderboard report
Generative 3D Object Classification Objaverse PointLLM-7B v1.2 Objaverse (Average) 53.00 #5 of 7 Archive leaderboard report
Generative 3D Object Classification Objaverse PointLLM-7B v1.2 Objaverse (C) 51.00 #5 of 7 Archive leaderboard report
Generative 3D Object Classification Objaverse PointLLM-7B v1.2 Objaverse (I) 55.00 #5 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections