Papers › Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

18 Mar 2024arXiv:2403.11401archive 2025-07-28

Rao Fu, Jingyu Liu, Xilun Chen, Yixin Nie, Wenhan Xiong

This paper introduces Scene-LLM, a 3D-visual-language model that enhances embodied agents' abilities in interactive 3D indoor environments by integrating the reasoning strengths of Large Language Models (LLMs). Scene-LLM adopts a hybrid 3D visual feature representation, that incorporates dense spatial information and supports scene state updates. The model employs a projection layer to efficiently project these features in the pre-trained textual embedding space, enabling effective interpretation of 3D visual information. Unique to our approach is the integration of both scene-level and ego-centric 3D information. This combination is pivotal for interactive planning, where scene-level data supports global planning and ego-centric data is important for localization. Notably, we use ego-centric 3D frame features for feature alignment, an efficient technique that enhances the model's ability to align features of small objects within the scene. Our experiments with Scene-LLM demonstrate its strong capabilities in dense captioning, question answering, and interactive planning. We believe Scene-LLM advances the field of 3D visual understanding and reasoning, offering new possibilities for sophisticated agent interactions in indoor settings.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Question Answering (3D-QA)Dense CaptioningLanguage ModelingLanguage ModellingQuestion Answering

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D Question Answering (3D-QA) SQA3D Scene-LLM Exact Match 54.2 #5 of 13 Archive leaderboard report
3D Question Answering (3D-QA) ScanQA Test w/ objects Scene-LLM BLEU-4 12.0 #4 of 18 Archive leaderboard report
3D Question Answering (3D-QA) ScanQA Test w/ objects Scene-LLM CIDEr 80 #4 of 18 Archive leaderboard report
3D Question Answering (3D-QA) ScanQA Test w/ objects Scene-LLM Exact Match 27.2 #4 of 18 Archive leaderboard report
3D Question Answering (3D-QA) ScanQA Test w/ objects Scene-LLM METEOR 16.6 #4 of 18 Archive leaderboard report
3D Question Answering (3D-QA) ScanQA Test w/ objects Scene-LLM ROUGE 40.0 #4 of 18 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

ALIGN

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections