Browse State-of-the-Art › Visual Grounding
Visual Grounding
299 papers with code · 4 benchmarks · 11 datasets archive 2025-07-28
Visual Grounding (VG) aims to locate the most relevant object or region in an image, based on a natural language query. The query can be a phrase, a sentence, or even a multi-round dialogue. There are three main challenges in VG:
- What is the main focus in a query?
- How to understand an image?
- How to locate an object?
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
4 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| RefCOCO+ testA (7 rows) | Florence-2-large-ft | Florence-2: Advancing a Unified Representation for a Variety of... | code | — | Compare |
| RefCOCO+ test B (6 rows) | Florence-2-large-ft | Florence-2: Advancing a Unified Representation for a Variety of... | code | — | Compare |
| RefCOCO+ val (6 rows) | Florence-2-large-ft | Florence-2: Advancing a Unified Representation for a Variety of... | code | — | Compare |
| RefCOCO testA (1 row) | HYDRA | HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning | code | Syntology ran 5 of 8 samples · 3 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
11 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
3 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 299 papers with code (571 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Aug 2019 11 repositories listed Syntology ran 10 of 34 samples · 24 unverified · 34 pointer-only (licence)We present ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language.
-
6 Jun 2016 10 repositories listedApproaches to multimodal pooling include element-wise product or sum, as well as concatenation of the visual and textual representations.
-
29 Mar 2023 7 repositories listed Syntology ran 8 of 12 samples · 4 unverified · 12 pointer-only (licence)In this paper, we propose ViewRefer, a multi-view framework for 3D visual grounding exploring how to grasp the view knowledge from both text and 3D modalities.
-
26 Apr 2021 5 repositories listed Syntology ran 6 of 11 samples · 5 unverifiedWe also investigate the utility of our model as an object detector on a given label set when fine-tuned in a few-shot setting.
-
28 Dec 2024 4 repositories listed Syntology ran 9 of 15 samples · 6 unverified · 3 pointer-only (licence)Finally, we outline the challenges confronting visual grounding and propose valuable directions for future research, which may serve as inspiration for subsequent researchers.
-
17 Oct 2023 4 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe present Set-of-Mark (SoM), a new visual prompting method, to unleash the visual grounding abilities of large multimodal models (LMMs), such as GPT-4V.
-
1 Feb 2023 4 repositories listed Syntology ran 9 of 19 samples · 10 unverifiedIn contrast to predominant paradigms of solely relying on sequence-to-sequence generation or encoder-based instance discrimination, mPLUG-2 introduces a multi-module composition network by sharing common universal…
-
7 Feb 2022 4 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn this work, we pursue a unified paradigm for multimodal pretraining to break the scaffolds of complex task/modality-specific customization.
-
23 May 2025 3 repositories listedTo address this limitation, we introduce OrionBench, a benchmark designed to support the development of accurate object detection models for charts and HROs in infographics.
-
18 Jun 2024 3 repositories listed Syntology ran 11 of 13 samples · 2 unverified · 13 pointer-only (licence)We introduce a new benchmark designed to advance the development of general-purpose, large-scale vision-language models for remote sensing images.
-
17 Jun 2024 3 repositories listed Syntology ran 1 of 5 samples · 4 unverified · 1 pointer-only (licence)Scientific documents record research findings and valuable human knowledge, comprising a vast corpus of high-quality data.
-
27 Feb 2024 3 repositories listed Syntology ran 9 of 17 samples · 8 unverifiedThis paper presents ShapeLLM, the first 3D Multimodal Large Language Model (LLM) designed for embodied interaction, exploring a universal 3D object understanding with 3D point clouds and languages.
-
15 May 2023 3 repositories listedIn order to utilize vision and language pre-trained models to address the grounding problem, and reasonably take advantage of pseudo-labels, we propose CLIP-VG, a novel method that can conduct self-paced curriculum…
-
29 Sep 2022 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)3D visual grounding aims to find the object within point clouds mentioned by free-form natural language descriptions with rich semantic cues.
-
24 May 2022 3 repositories listedLarge-scale pretrained foundation models have been an emerging paradigm for building artificial intelligence (AI) systems, which can be quickly adapted to a wide range of downstream tasks.
-
30 Mar 2022 3 repositories listed Syntology ran 5 of 7 samples · 2 unverifiedTo implement this idea, we propose Collaborative Glance-Gaze TransFormer (CoFormer) that consists of two modules: Glance transformer for activity classification and Gaze transformer for entity estimation.
-
30 Mar 2022 3 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedIn this paper, we propose a simple yet universal network termed SeqTR for visual grounding tasks, e.
-
28 Mar 2022 3 repositories listedWe present a method for visually-grounded spoken term discovery.
-
10 Sep 2018 3 repositories listedWe compare our approach to an alternative system which extends the baseline with reinforcement learning.
-
27 Jun 2016 3 repositories listedVisual question answering (VQA) is an interesting learning setting for evaluating the abilities and shortcomings of current systems for image understanding.
-
12 Nov 2015 3 repositories listedWe propose a novel approach which learns grounding by reconstructing a given phrase using an attention mechanism, which can be either latent or optimized directly.
-
18 May 2025 2 repositories listedGraphical User Interface (GUI) agents have made substantial strides in understanding and executing user instructions across diverse platforms.
-
21 Jan 2025 2 repositories listedTo address this, we identify two critical sets of visual tokens that facilitate the transfer of visual information from the vision encoder to the LLM.
-
25 Nov 2024 2 repositories listedAdvances in multimodal pre-training have propelled object-level foundation models, such as Grounding DINO and Florence-2, in tasks like visual grounding and object detection.
-
11 Jun 2024 2 repositories listedGrounded Multimodal Named Entity Recognition (GMNER) task aims to identify named entities, entity types and their corresponding visual regions.
-
9 Jun 2024 2 repositories listedTo address this issue, we present F-LMM -- grounding frozen off-the-shelf LMMs in human-AI conversations -- a straightforward yet effective design based on the fact that word-pixel correspondences conducive to visual…
-
29 Mar 2024 2 repositories listed Syntology ran 2 of 4 samples · 2 unverifiedVHM is built on a large-scale remote sensing image-text dataset with rich-content captions (VersaD), and an honest instruction dataset comprising both factual and deceptive questions (HnstD).
-
21 Mar 2024 2 repositories listedToday's most accurate language models are trained on orders of magnitude more language data than human language learners receive - but with no supervision from other sensory modalities that play a crucial role in human…
-
23 Feb 2024 2 repositories listed Syntology ran 11 of 17 samples · 6 unverified · 17 pointer-only (licence)Large Vision-Language Models (LVLMs) are susceptible to object hallucinations, an issue in which their generated text contains non-existent objects, greatly limiting their reliability and practicality.
-
15 Feb 2024 2 repositories listedGrounded Multimodal Named Entity Recognition (GMNER) is a nascent multimodal task that aims to identify named entities, entity types and their corresponding visual regions.
Syntology lines on 15 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections