Browse State-of-the-Art › 3D visual grounding
3D visual grounding
39 papers with code · 0 benchmarks · 2 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 39 papers with code (82 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
29 Mar 2023 7 repositories listed Syntology ran 8 of 12 samples · 4 unverified · 12 pointer-only (licence)In this paper, we propose ViewRefer, a multi-view framework for 3D visual grounding exploring how to grasp the view knowledge from both text and 3D modalities.
-
29 Sep 2022 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)3D visual grounding aims to find the object within point clouds mentioned by free-form natural language descriptions with rich semantic cues.
-
14 Mar 2021 2 repositories listed Syntology ran 11 of 16 samples · 5 unverified · 6 pointer-only (licence)Grounding referring expressions in RGBD image has been an emerging field.
-
13 May 2025 1 repository listedFurthermore, we propose DriveMonkey, a simple yet effective framework that seamlessly integrates LVLMs with a spatial processor using a series of learnable queries.
-
7 May 2025 1 repository listed3D visual grounding aims to localize the unique target described by natural languages in 3D scenes.
-
13 Apr 2025 1 repository listed3-Dimensional Embodied Reference Understanding (3D-ERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene.
-
28 Mar 2025 1 repository listed Syntology ran 1 of 8 samples · 7 unverifiedOur evaluation of state-of-the-art 3D-VL models on Beacon3D reveals that (i) object-centric evaluation elicits true model performance and particularly weak generalization in QA; (ii) grounding-QA coherence remains…
-
14 Feb 2025 1 repository listedTo this end, we propose text-guided pruning (TGP) and completion-based addition (CBA) to deeply fuse 3D scene representation and text features in an efficient way by gradual region pruning and target completion.
-
3 Feb 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverified3D visual grounding (3DVG) is challenging because of the requirement of understanding on visual information, language and spatial relationships.
-
1 Jan 2025 1 repository listed3-Dimensional Embodied Reference Understanding (3DERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene.
-
22 Nov 2024 1 repository listedIn embodied intelligence systems, a key component is 3D perception algorithm, which enables agents to understand their surrounding environments.
-
21 Nov 2024 1 repository listed3D visual grounding (3DVG) aims to locate objects in a 3D scene with natural language descriptions.
-
17 Oct 2024 1 repository listed Syntology ran 10 of 11 samples · 1 unverified · 11 pointer-only (licence)3D visual grounding is crucial for robots, requiring integration of natural language and 3D scene understanding.
-
25 Jul 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)3D referring segmentation is an emerging and challenging vision-language task that aims to segment the object described by a natural language expression in a point cloud scene.
-
7 Jul 2024 1 repository listed3D referring expression comprehension (3DREC) and segmentation (3DRES) have overlapping objectives, indicating their potential for collaboration.
-
13 Jun 2024 1 repository listedWith the emergence of LLMs and their integration with other data modalities, multi-modal 3D perception attracts more attention due to its connectivity to the physical world and makes rapid progress.
-
9 Jun 2024 1 repository listedIn this survey, we attempt to provide a comprehensive overview of the T-3DVG progress, including its fundamental elements, recent research advances, and future research directions.
-
Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression Comprehension21 May 2024 1 repository listedMoreover, we propose a novel model, T-RadarNet, for 3D REC on point clouds, achieving State-Of-The-Art (SOTA) performance on the Talk2Radar dataset compared to counterparts.
-
17 Apr 2024 1 repository listed3D Visual Grounding (3DVG) and 3D Dense Captioning (3DDC) are two crucial tasks in various 3D applications, which require both shared and complementary information in localization and visual-language relationships.
-
13 Mar 2024 1 repository listed3D visual grounding aims to automatically locate the 3D region of the specified object given the corresponding textual description.
-
5 Mar 2024 1 repository listed Syntology ran 11 of 16 samples · 5 unverified · 16 pointer-only (licence)3D visual grounding involves matching natural language descriptions with their corresponding objects in 3D spaces.
-
1 Jan 2024 1 repository listed3D visual grounding aims to localize 3D objects described by free-form language sentences.
-
13 Dec 2023 1 repository listed Syntology ran 5 of 7 samples · 2 unverified · 7 pointer-only (licence)To foster this task, we propose Mono3DVG-TR, an end-to-end transformer-based network, which takes advantage of both the appearance and geometry information in text embeddings for multi-modal learning and 3D object…
-
26 Nov 2023 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Building on this, we design a visual program that consists of three types of modules, i.
-
28 Oct 2023 1 repository listed Syntology ran 9 of 11 samples · 2 unverifiedTo tackle this problem, we introduce the CityRefer dataset for city-level visual grounding.
-
10 Oct 2023 1 repository listed Syntology ran 10 of 11 samples · 1 unverified · 11 pointer-only (licence)To this end, we formulate the 3D visual grounding problem as a sequence-to-sequence Seq2Seq task by first predicting a chain of anchors and then the final target.
-
21 Sep 2023 1 repository listed Syntology ran 5 of 5 samples · 0 unverifiedWhile existing approaches often rely on extensive labeled data or exhibit limitations in handling complex language queries, we propose LLM-Grounder, a novel zero-shot, open-vocabulary, Large Language Model (LLM)-based…
-
11 Sep 2023 1 repository listed Syntology ran 5 of 8 samples · 3 unverifiedWe introduce the task of localizing a flexible number of objects in real-world 3D scenes using natural language descriptions.
-
18 Jul 2023 1 repository listedTo accomplish this, we design a novel semantic matching model that analyzes the semantic similarity between object proposals and sentences in a coarse-to-fine manner.
-
23 May 2023 1 repository listedWe present a novel task for cross-dataset visual grounding in 3D scenes (Cross3DVG), which overcomes limitations of existing 3D visual grounding models, specifically their restricted 3D resources and consequent…
Syntology lines on 14 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections