Browse State-of-the-Art › Phrase Grounding
Phrase Grounding
51 papers with code · 5 benchmarks · 6 datasets archive 2025-07-28
Given an image and a corresponding caption, the Phrase Grounding task aims to ground each entity mentioned by a noun phrase in the caption to a region in the image.
Source: Phrase Grounding by Soft-Label Chain Conditional Random Field
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Flickr30k Entities Test (18 rows) | GLIPv2 | GLIPv2: Unifying Localization and Vision-Language Understanding | code | Syntology ran 2 of 2 samples · 0 unverified | Compare |
| Flickr30k (3 rows) | GBS Ensemble + 12-in-1 | Detector-Free Weakly Supervised Grounding by Separation | code | — | Compare |
| Flickr30k Entities Dev (3 rows) | Fiber-B | Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone | code | Syntology ran 2 of 3 samples · 1 unverified | Compare |
| ReferIt (3 rows) | VG_BiLSTM_VGG | Multi-level Multimodal Common Semantic Space for Image-Phrase Grounding | code | — | Compare |
| Visual Genome (3 rows) | GbS VG | Detector-Free Weakly Supervised Grounding by Separation | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 51 papers with code (88 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Jun 2016 10 repositories listedApproaches to multimodal pooling include element-wise product or sum, as well as concatenation of the visual and textual representations.
-
26 Apr 2021 5 repositories listed Syntology ran 6 of 11 samples · 5 unverifiedWe also investigate the utility of our model as an object detector on a given label set when fine-tuned in a few-shot setting.
-
28 Dec 2024 4 repositories listed Syntology ran 9 of 15 samples · 6 unverified · 3 pointer-only (licence)Finally, we outline the challenges confronting visual grounding and propose valuable directions for future research, which may serve as inspiration for subsequent researchers.
-
17 Nov 2018 3 repositories listedMost existing work that grounds natural language phrases in images starts with the assumption that the phrase in question is relevant to the image.
-
12 Nov 2015 3 repositories listedWe propose a novel approach which learns grounding by reconstructing a given phrase using an attention mechanism, which can be either latent or optimized directly.
-
4 Jan 2024 2 repositories listed Syntology ran 4 of 6 samples · 2 unverifiedGrounding-DINO is a state-of-the-art open-set detection model that tackles multiple vision tasks including Open-Vocabulary Detection (OVD), Phrase Grounding (PG), and Referring Expression Comprehension (REC).
-
26 Jun 2023 2 repositories listedWe introduce Kosmos-2, a Multimodal Large Language Model (MLLM), enabling new capabilities of perceiving object descriptions (e.
-
21 Apr 2022 2 repositories listed Syntology ran 3 of 4 samples · 1 unverifiedWe release a new dataset with locally-aligned phrase grounding annotations by radiologists to facilitate the study of complex semantic modelling in biomedical vision--language processing.
-
19 May 2015 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)The Flickr30k dataset has become a standard benchmark for sentence-based image description.
-
16 May 2025 1 repository listed Syntology ran 2 of 6 samples · 4 unverifiedPhrase grounding between images and their captions is a well-established task.
-
23 Feb 2025 1 repository listedMedical Phrase Grounding (MPG) maps radiological findings described in medical reports to specific regions in medical images.
-
29 Jan 2025 1 repository listedAs artificial intelligence (AI) becomes increasingly central to healthcare, the demand for explainable and trustworthy models is paramount.
-
16 Oct 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Many artwork collections contain textual attributes that provide rich and contextualised descriptions of artworks.
-
13 Sep 2024 1 repository listedIn this paper, we address a challenging task, synchronous motion captioning, that aim to generate a language description synchronized with human motion sequences.
-
20 Aug 2024 1 repository listedThe core of our approach involves freezing the MDETR backbone and training only the Universal Projection module (UP), which bridges vision and language representations.
-
1 Jul 2024 1 repository listedWe introduce the concept of "empathic grounding" in conversational agents as an extension of Clark's conceptualization of grounding in conversation in which the grounding criterion includes listener empathy for the…
-
19 Apr 2024 1 repository listed Syntology ran 7 of 9 samples · 2 unverifiedIn addition, we aim to perform this task in a zero-shot manner, i.
-
14 Mar 2024 1 repository listedLarge Vision Language Models have achieved fine-grained object perception, but the limitation of image resolution remains a significant obstacle to surpass the performance of task-specific experts in complex and dense…
-
22 Nov 2023 1 repository listed Syntology ran 1 of 5 samples · 4 unverified · 5 pointer-only (licence)Extending image-based Large Multimodal Models (LMMs) to videos is challenging due to the inherent complexity of video data.
-
5 Nov 2023 1 repository listedWhile we demonstrate our data augmentation method with MDETR framework, the proposed approach is applicable to common grounding-based vision and language tasks with other frameworks.
-
23 Oct 2023 1 repository listedThe ability to actively ground task instructions from an egocentric view is crucial for AI agents to accomplish tasks or assist humans virtually.
-
12 Sep 2023 1 repository listedIt designs a correlation weighting mechanism to adjust the correlation between masked chest X-ray image patches and their corresponding reports, thereby enhancing the model's representation learning capabilities.
-
7 Sep 2023 1 repository listed Syntology ran 6 of 11 samples · 5 unverified · 11 pointer-only (licence)It has been established that training a box-based detector network can enhance the localization performance of weakly supervised and unsupervised methods.
-
6 Sep 2023 1 repository listed Syntology ran 10 of 12 samples · 2 unverified · 12 pointer-only (licence)Key to tasks that require reasoning about natural language in visual contexts is grounding words and phrases to image regions.
-
5 Sep 2023 1 repository listedIn recent years, cross-modal reasoning (CMR), the process of understanding and reasoning across different modalities, has emerged as a pivotal area with applications spanning from multimedia analysis to healthcare…
-
31 Mar 2023 1 repository listedRecent advancements in diffusion models have significantly impacted the trajectory of generative machine learning research, with many adopting the strategy of fine-tuning pre-trained models using domain-specific…
-
11 Jan 2023 1 repository listedPrior work in biomedical VLP has mostly relied on the alignment of single image and report pairs even though clinical notes commonly refer to prior images.
-
1 Jan 2023 1 repository listedA phrase grounding model receives an input image and a text phrase and outputs a suitable localization map.
-
28 Nov 2022 1 repository listedAs phrase extraction can be regarded as a $1$D text segmentation problem, we formulate PEG as a dual detection problem and propose a novel DQ-DETR model, which introduces dual queries to probe different features from…
-
23 Oct 2022 1 repository listed Syntology ran 0 of 9 samples · 9 unverifiedFirst, we construct a dataset of phrase grounding with both noun phrases and pronouns to image regions.
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections