{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/3dgraphllm-combining-semantic-graphs-and","title":"3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding","arxiv_id":"2412.18450","date":"2024-12-24","proceeding":null,"authors":["Tatiana Zemskova","Dmitry Yudin"],"abstract":"A 3D scene graph represents a compact scene model, storing information about the objects and the semantic relationships between them, making its use promising for robotic tasks. When interacting with a user, an embodied intelligent agent should be capable of responding to various queries about the scene formulated in natural language. Large Language Models (LLMs) are beneficial solutions for user-robot interaction due to their natural language understanding and reasoning abilities. Recent methods for creating learnable representations of 3D scenes have demonstrated the potential to improve the quality of LLMs responses by adapting to the 3D world. However, the existing methods do not explicitly utilize information about the semantic relationships between objects, limiting themselves to information about their coordinates. In this work, we propose a method 3DGraphLLM for constructing a learnable representation of a 3D scene graph. The learnable representation is used as input for LLMs to perform 3D vision-language tasks. In our experiments on popular ScanRefer, RIORefer, Multi3DRefer, ScanQA, Sqa3D, and Scan2cap datasets, we demonstrate the advantage of this approach over baseline methods that do not use information about the semantic relationships between objects. The code is publicly available at https://github.com/CognitiveAISystems/3DGraphLLM.","url_abs":"https://arxiv.org/abs/2412.18450v2","url_pdf":"https://arxiv.org/pdf/2412.18450v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"3dgraphllm-combining-semantic-graphs-and","repo_url":"https://github.com/cognitiveaisystems/3dgraphllm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"natural-language-understanding","task_name":"Natural Language Understanding"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2412.18450","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2412.18450"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/cognitiveaisystems/3dgraphllm","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_fixture":1,"ran_draft_wrong":3,"ran":4},"by_repo_kind":{"official":{"samples":8,"ran":8,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"30d7eec482ebf6b1","entry":"repeat_kv","repo":"cognitiveaisystems/3dgraphllm","repo_kind":"official","path":"models/modeling_llama.py","file_url":"https://github.com/cognitiveaisystems/3dgraphllm/blob/HEAD/models/modeling_llama.py","link_basis":"harvester_set","language":"python","status":"ran_fixture","verification_level":2,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"30d7eec482ebf6b1"}},{"code_sha256_prefix":"bac65c3dafaec040","entry":"apply_rotary_pos_emb","repo":"cognitiveaisystems/3dgraphllm","repo_kind":"official","path":"models/modeling_llama.py","file_url":"https://github.com/cognitiveaisystems/3dgraphllm/blob/HEAD/models/modeling_llama.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bac65c3dafaec040"}},{"code_sha256_prefix":"cda0cda020ca6dee","entry":"get_caption","repo":"cognitiveaisystems/3dgraphllm","repo_kind":"official","path":"preprocess/prepare_scannet_region_caption_annos.py","file_url":"https://github.com/cognitiveaisystems/3dgraphllm/blob/HEAD/preprocess/prepare_scannet_region_caption_annos.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"cda0cda020ca6dee"}},{"code_sha256_prefix":"891b8ebab395921f","entry":"get_clones","repo":"cognitiveaisystems/3dgraphllm","repo_kind":"official","path":"models/helpers.py","file_url":"https://github.com/cognitiveaisystems/3dgraphllm/blob/HEAD/models/helpers.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"891b8ebab395921f"}},{"code_sha256_prefix":"3f8b2fc0609a5184","entry":"nclamp","repo":"cognitiveaisystems/3dgraphllm","repo_kind":"official","path":"models/graph3dllm.py","file_url":"https://github.com/cognitiveaisystems/3dgraphllm/blob/HEAD/models/graph3dllm.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3f8b2fc0609a5184"}},{"code_sha256_prefix":"3895c5e4b8839cf6","entry":"recover_caption","repo":"cognitiveaisystems/3dgraphllm","repo_kind":"official","path":"preprocess/prepare_scannet_region_caption_annos.py","file_url":"https://github.com/cognitiveaisystems/3dgraphllm/blob/HEAD/preprocess/prepare_scannet_region_caption_annos.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3895c5e4b8839cf6"}},{"code_sha256_prefix":"b99eea6376d1e212","entry":"rotate_half","repo":"cognitiveaisystems/3dgraphllm","repo_kind":"official","path":"models/modeling_llama.py","file_url":"https://github.com/cognitiveaisystems/3dgraphllm/blob/HEAD/models/modeling_llama.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b99eea6376d1e212"}},{"code_sha256_prefix":"8ea16420707b36f4","entry":"update_caption","repo":"cognitiveaisystems/3dgraphllm","repo_kind":"official","path":"preprocess/prepare_scannet_region_caption_annos.py","file_url":"https://github.com/cognitiveaisystems/3dgraphllm/blob/HEAD/preprocess/prepare_scannet_region_caption_annos.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8ea16420707b36f4"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}