Papers › ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations

24 Jan 2025arXiv:2501.14607archive 2025-07-28

Tianming Liang, Kun-Yu Lin, Chaolei Tan, JianGuo Zhang, Wei-Shi Zheng, Jian-Fang Hu

Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. Despite notable progress in recent years, current RVOS models remain struggle to handle complicated object descriptions due to their limited video-language understanding. To address this limitation, we present \textbf{ReferDINO}, an end-to-end RVOS model that inherits strong vision-language understanding from the pretrained visual grounding foundation models, and is further endowed with effective temporal understanding and object segmentation capabilities. In ReferDINO, we contribute three technical innovations for effectively adapting the foundation models to RVOS: 1) an object-consistent temporal enhancer that capitalizes on the pretrained object-text representations to enhance temporal understanding and object consistency; 2) a grounding-guided deformable mask decoder that integrates text and grounding conditions to generate accurate object masks; 3) a confidence-aware query pruning strategy that significantly improves the object decoding efficiency without compromising performance. We conduct extensive experiments on five public RVOS benchmarks to demonstrate that our proposed ReferDINO outperforms state-of-the-art methods significantly. Project page: \url{https://isee-laboratory.github.io/ReferDINO}

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderObjectReferring Expression SegmentationReferring Video Object SegmentationSemantic SegmentationVideo Object SegmentationVideo Semantic SegmentationVisual Grounding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) ReferDINO (Swin-B) F 71.5 #5 of 33 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) ReferDINO (Swin-B) J 67.0 #5 of 33 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) ReferDINO (Swin-B) J&F 69.3 #5 of 33 Archive leaderboard report
Referring Video Object Segmentation Long-RVOS ReferDINO J&F 48.7 #2 of 7 Archive leaderboard report
Referring Video Object Segmentation Long-RVOS ReferDINO tIoU 71.7 #2 of 7 Archive leaderboard report
Referring Video Object Segmentation Long-RVOS ReferDINO vIoU 41.2 #2 of 7 Archive leaderboard report
Referring Video Object Segmentation MeViS ReferDINO (Swin-B) F 53.9 #5 of 16 Archive leaderboard report
Referring Video Object Segmentation MeViS ReferDINO (Swin-B) J 44.7 #5 of 16 Archive leaderboard report
Referring Video Object Segmentation MeViS ReferDINO (Swin-B) J&F 49.3 #5 of 16 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Pruning

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections