Browse State-of-the-Art › Referring Video Object Segmentation
Referring Video Object Segmentation
50 papers with code · 5 benchmarks · 4 datasets archive 2025-07-28
Referring video object segmentation aims at segmenting an object in video with language expressions. Unlike the previous video object segmentation, the task exploits a different type of supervision, language expressions, to identify and segment an object referred by the given language expressions in a video.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Refer-YouTube-VOS (18 rows) | FindTrack | Find First, Track Next: Decoupling Identification and Propagation... | code | — | Compare |
| MeViS (16 rows) | MPG-SAM 2 | MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for... | code | Syntology ran 5 of 16 samples · 11 unverified | Compare |
| Ref-DAVIS17 (11 rows) | FindTrack | Find First, Track Next: Decoupling Identification and Propagation... | code | — | Compare |
| ReVOS (9 rows) | VRS-HQ (Chat-UniVi-13B) | The Devil is in Temporal Token: High Quality Video Reasoning Segmentation | code | — | Compare |
| Long-RVOS (7 rows) | ReferMo | Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video... | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 50 papers with code (74 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
16 Jul 2024 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)In this paper, we introduce a new task, Reasoning Video Object Segmentation (ReasonVOS).
-
25 Dec 2023 2 repositories listed Syntology ran 9 of 9 samples · 0 unverifiedWe evaluate our unified models on various benchmarks.
-
1 Aug 2023 2 repositories listedIn this work, we propose a new segmentation task -- reasoning segmentation.
-
29 Nov 2021 2 repositories listed Syntology ran 6 of 11 samples · 5 unverifiedDue to the complex nature of this multimodal task, which combines text reasoning, video understanding, instance segmentation and tracking, existing approaches typically rely on sophisticated pipelines in order to tackle…
-
5 Jun 2025 1 repository listedCurrent video-based approaches, while proficient in tracking, lack the sophisticated reasoning capabilities of large language models, limiting their contextual understanding and generalization.
-
18 Apr 2025 1 repository listedReferring video object segmentation (RVOS) aims to segment objects in videos guided by natural language descriptions.
-
10 Apr 2025 1 repository listed Syntology ran 3 of 12 samples · 9 unverified · 12 pointer-only (licence)This paper proposes a novel framework utilizing multi-modal large language models (MLLMs) for referring video object segmentation (RefVOS).
-
7 Apr 2025 1 repository listedMotion expression video segmentation is designed to segment objects in accordance with the input motion expressions.
-
1 Apr 2025 1 repository listedReferring video object segmentation (RVOS) is a challenging task that requires the model to segment the object in a video given the language description.
-
30 Mar 2025 1 repository listedThis task has attracted increasing attention in the field of computer vision due to its promising applications in video editing and human-agent interaction.
-
5 Mar 2025 1 repository listedThis reference is then utilized by a dedicated propagation module to track and segment the object across the entire video.
-
23 Jan 2025 1 repository listed Syntology ran 5 of 16 samples · 11 unverifiedReferring video object segmentation (RVOS) aims to segment objects in a video according to textual descriptions, which requires the integration of multimodal information and temporal dynamics perception.
-
21 Jan 2025 1 repository listedThis paper aims to improve the performance of video multimodal large language models (MLLM) via long and rich context (LRC) modeling.
-
15 Jan 2025 1 repository listedExisting methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial complexity and inter-frame motion.
-
9 Jan 2025 1 repository listedThe Aligner removes noise from queries and aligns them to achieve query consistency.
-
7 Jan 2025 1 repository listedThis work presents Sa2VA, the first unified model for dense grounded understanding of both images and videos.
-
1 Jan 2025 1 repository listedWe identify three critical challenges: (C1) insufficient quantitative representation of textual numerical data, (C2) repetitive and degraded response templates for spatiotemporal referencing, and (C3) loss of visual…
-
2 Dec 2024 1 repository listedReferring video object segmentation (RVOS) requires tracking and segmenting an object throughout a video according to a given natural language expression, demanding both complex motion understanding and the alignment of…
-
26 Nov 2024 1 repository listedReferring Video Object Segmentation (RVOS) relies on natural language expressions to segment an object in a video clip.
-
26 Nov 2024 1 repository listed Syntology ran 7 of 17 samples · 10 unverifiedThis paper aims to address universal segmentation for image and video perception with the strong reasoning ability empowered by Visual Large Language Models (VLLMs).
-
29 Sep 2024 1 repository listed Syntology ran 7 of 16 samples · 9 unverifiedWe introduce VideoLISA, a video-based multimodal large language model designed to tackle the problem of language-instructed reasoning segmentation in videos.
-
10 Jul 2024 1 repository listed Syntology ran 14 of 15 samples · 1 unverified · 15 pointer-only (licence)Experimental results on VISOR dataset reveal that ActionVOS significantly reduces the mis-segmentation of inactive objects, confirming that actions help the ActionVOS model understand objects' involvement.
-
11 Jun 2024 1 repository listedMotion Expression guided Video Segmentation (MeViS), as an emerging task, poses many new challenges to the field of referring video object segmentation (RVOS).
-
12 Apr 2024 1 repository listedThe complexity of this task increases with the intricacy of the sentences provided.
-
4 Apr 2024 1 repository listed Syntology ran 2 of 7 samples · 5 unverified · 7 pointer-only (licence)In fact, static cues can sometimes interfere with temporal perception by overshadowing motion cues.
-
28 Mar 2024 1 repository listed Syntology ran 14 of 15 samples · 1 unverifiedReferring Video Object Segmentation (R-VOS) methods face challenges in maintaining consistent object segmentation due to temporal context variability and the presence of other visually similar objects.
-
18 Mar 2024 1 repository listed Syntology ran 16 of 20 samples · 4 unverified · 20 pointer-only (licence)We hypothesize that the latent representation learned from a pretrained generative T2V model encapsulates rich semantics and coherent temporal correspondences, thereby naturally facilitating video understanding.
-
28 Feb 2024 1 repository listed Syntology ran 12 of 14 samples · 2 unverified · 14 pointer-only (licence)Despite the recent advances in unified image segmentation (IS), developing a unified video segmentation (VS) model remains a challenge.
-
1 Jan 2024 1 repository listedThe recent transformer-based models have dominated the Referring Video Object Segmentation (RVOS) task due to the superior performance.
-
29 Dec 2023 1 repository listedThe perception component then generates the tracking results based on the embeddings.
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections