Papers › LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
Chenming Zhu, Tai Wang, Wenwei Zhang, Jiangmiao Pang, Xihui Liu
Recent advancements in Large Multimodal Models (LMMs) have greatly enhanced their proficiency in 2D visual understanding tasks, enabling them to effectively process and understand images and videos. However, the development of LMMs with 3D-awareness for 3D scene understanding has been hindered by the lack of large-scale 3D vision-language datasets and powerful 3D encoders. In this paper, we introduce a simple yet effective framework called LLaVA-3D. Leveraging the strong 2D understanding priors from LLaVA, our LLaVA-3D efficiently adapts LLaVA for 3D scene understanding without compromising 2D understanding capabilities. To achieve this, we utilize the 3D position embeddings to bring the 2D CLIP patches within a 3D spatial context. By integrating the 3D position embeddings into 2D LMMs and employing joint 2D and 3D vision-language instruction tuning, we establish a unified architecture for both 2D image understanding and 3D scene understanding. Experimental results show that LLaVA-3D converges 3.5x faster than existing 3D LMMs when trained on 3D vision-language datasets. Moreover, LLaVA-3D not only achieves state-of-the-art performance across various 3D tasks but also maintains comparable 2D image understanding and vision-language conversation capabilities with LLaVA.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| 3D Question Answering (3D-QA) | SQA3D | LLaVA-3D | Exact Match | 60.1 | #1 of 13 | Archive leaderboard | report |
| 3D Question Answering (3D-QA) | ScanQA Test w/ objects | LLaVA-3D | BLEU-4 | 16.4 | #2 of 18 | Archive leaderboard | report |
| 3D Question Answering (3D-QA) | ScanQA Test w/ objects | LLaVA-3D | CIDEr | 103.1 | #2 of 18 | Archive leaderboard | report |
| 3D Question Answering (3D-QA) | ScanQA Test w/ objects | LLaVA-3D | Exact Match | 30.6 | #2 of 18 | Archive leaderboard | report |
| 3D Question Answering (3D-QA) | ScanQA Test w/ objects | LLaVA-3D | METEOR | 20.8 | #2 of 18 | Archive leaderboard | report |
| 3D Question Answering (3D-QA) | ScanQA Test w/ objects | LLaVA-3D | ROUGE | 49.6 | #2 of 18 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections