Browse State-of-the-Art › Referring Expression Segmentation
Referring Expression Segmentation
97 papers with code · 22 benchmarks · 11 datasets archive 2025-07-28
The task aims at labeling the pixels of an image or video that represent an object instance referred by a linguistic expression. In particular, the referring expression (RE) must allow the identification of an individual object in a discourse or scene (the referent). REs unambiguously identify the target instance.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
22 leaderboard tables shown for this task, 22 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 22 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
11 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 97 papers with code (145 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
18 Dec 2021 6 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)After training on an extended version of the PhraseCut dataset, our system generates a binary segmentation map for an image based on a free-text prompt or on an additional image expressing the query.
-
26 Apr 2021 5 repositories listed Syntology ran 6 of 11 samples · 5 unverifiedWe also investigate the utility of our model as an object detector on a given label set when fine-tuned in a few-shot setting.
-
20 Mar 2016 4 repositories listedTo produce pixelwise segmentation for the language expression, we propose an end-to-end trainable recurrent and convolutional network model that jointly learns to process visual and linguistic information.
-
17 May 2025 3 repositories listed Syntology ran 1 of 15 samples · 14 unverifiedLarge vision-language models exhibit inherent capabilities to handle diverse visual perception tasks.
-
9 Mar 2025 3 repositories listedTraditional methods for reasoning segmentation rely on supervised fine-tuning with categorical labels and simple descriptions, limiting its out-of-domain generalization and lacking explicit reasoning processes.
-
30 Mar 2022 3 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedIn this paper, we propose a simple yet universal network termed SeqTR for visual grounding tasks, e.
-
3 Jan 2019 3 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)Yet there has been evidence that current benchmark datasets suffer from bias, and current state-of-the-art models cannot be easily evaluated on their intermediate reasoning process.
-
30 Jul 2024 2 repositories listed Syntology ran 4 of 9 samples · 5 unverified · 9 pointer-only (licence)3D Referring Expression Segmentation (3D-RES) is dedicated to segmenting a specific instance within a 3D space based on a natural language description.
-
9 Jun 2024 2 repositories listedTo address this issue, we present F-LMM -- grounding frozen off-the-shelf LMMs in human-AI conversations -- a straightforward yet effective design based on the fact that word-pixel correspondences conducive to visual…
-
24 May 2024 2 repositories listedBy decoupling the intricate referring semantics into different granularity with a visual-linguistic hierarchy, and dynamic aggregating it with intra- and inter-selection, CoHD boosts multi-granularity comprehension with…
-
25 Dec 2023 2 repositories listed Syntology ran 9 of 9 samples · 0 unverifiedWe evaluate our unified models on various benchmarks.
-
4 Dec 2023 2 repositories listed Syntology ran 13 of 16 samples · 3 unverifiedThis paper aims to achieve universal segmentation of arbitrary semantic level.
-
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond24 Aug 2023 2 repositories listed Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)In this work, we introduce the Qwen-VL series, a set of large-scale vision-language models (LVLMs) designed to perceive and understand both texts and images.
-
1 Jun 2023 2 repositories listed Syntology ran 2 of 5 samples · 3 unverifiedExisting classic RES datasets and methods commonly support single-target expressions only, i.
-
3 Mar 2023 2 repositories listed Syntology ran 4 of 5 samples · 1 unverifiedIn this paper, we propose VPD (Visual Perception with a pre-trained Diffusion model), a new framework that exploits the semantic information of a pre-trained text-to-image diffusion model in visual perception tasks.
-
29 Nov 2021 2 repositories listed Syntology ran 6 of 11 samples · 5 unverifiedDue to the complex nature of this multimodal task, which combines text reasoning, video understanding, instance segmentation and tracking, existing approaches typically rely on sophisticated pipelines in order to tackle…
-
8 Jun 2021 2 repositories listedRecent advances in deep learning have brought significant progress in visual grounding tasks such as language-guided video object segmentation.
-
1 Oct 2020 2 repositories listedThe task of video object segmentation with referring expressions (language-guided VOS) is to, given a linguistic phrase and a video, generate binary masks for the object to which the phrase refers.
-
19 Mar 2020 2 repositories listed Syntology ran 0 of 13 samples · 13 unverifiedIn addition, we address a key challenge in this multi-task setup, i.
-
2 Jul 2025 1 repository listedReferring Image Segmentation (RIS) is a challenging task that aims to segment objects in an image based on natural language expressions.
-
23 May 2025 1 repository listedCombining these innovations on data and modeling, we present RemoteSAM, a foundation model that establishes new SoTA on several earth observation perception benchmarks, outperforming other foundation models such as…
-
13 Mar 2025 1 repository listedPixel grounding, encompassing tasks such as Referring Expression Segmentation (RES), has garnered considerable attention due to its immense potential for bridging the gap between vision and language modalities.
-
11 Mar 2025 1 repository listedWhile MLLMs have demonstrated adequate image understanding capabilities, they still struggle with pixel-level comprehension, limiting their practical applications.
-
6 Feb 2025 1 repository listedIn this work, we propose two novel challenging benchmarks and show that MLLMs without pixel-level grounding supervision can outperform the state of the art in such tasks when evaluating both the pixel-level grounding…
-
23 Jan 2025 1 repository listed Syntology ran 5 of 16 samples · 11 unverifiedReferring video object segmentation (RVOS) aims to segment objects in a video according to textual descriptions, which requires the integration of multimodal information and temporal dynamics perception.
-
21 Jan 2025 1 repository listedThis paper aims to improve the performance of video multimodal large language models (MLLM) via long and rich context (LRC) modeling.
-
15 Jan 2025 1 repository listedExisting methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial complexity and inter-frame motion.
-
15 Jan 2025 1 repository listed Syntology ran 7 of 17 samples · 10 unverifiedWhile some pioneering studies have undertaken preliminary explorations, they still remain at the level of aligned encoders (e.
-
12 Jan 2025 1 repository listed Syntology ran 5 of 13 samples · 8 unverifiedFurthermore, to address the challenge of insufficient multimodal understanding, we leverage pre-trained models based on visual-linguistic fusion representations.
-
9 Jan 2025 1 repository listedTo tackle intent ambiguity, we designed a Prompt-Aware Decoder (PAD) that guides the decoding process by deriving task-driven signals from the interaction between the expression and visual features.
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections