Papers › URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark

URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark

1 Aug 2020ECCV 2020 8archive 2025-07-28

Seonguk Seo, Joon-Young Lee, Bohyung Han

We propose a unified referring video object segmentation network (URVOS). URVOS takes a video and a referring expression as inputs, and estimates the {object masks} referred by the given language expression in the whole video frames. Our algorithm addresses the challenging problem by performing language-based object segmentation and mask propagation jointly using a single deep neural network with a proper combination of two attention models. In addition, we construct the first large-scale referring video object segmentation dataset called Refer-Youtube-VOS. We evaluate our model on two benchmark datasets including ours and demonstrate the effectiveness of the proposed approach. The dataset is released at \url{https://github.com/skynbe/Refer-Youtube-VOS}.

PaperPDFCode

Code

skynbe/Refer-Youtube-VOS officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ObjectOne-shot visual object segmentationReferring ExpressionReferring Expression SegmentationReferring Video Object SegmentationSegmentationSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentation

Datasets

Introduced by this paper, per the archive.

Refer-YouTube-VOS

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation DAVIS 2017 (val) URVOS + Refer-Youtube-VOS + ft. DAVIS J&F 1st frame 51.63 #9 of 18 Archive leaderboard report
Referring Expression Segmentation DAVIS 2017 (val) URVOS + Refer-Youtube-VOS J&F 1st frame 46.85 #11 of 18 Archive leaderboard report
Referring Expression Segmentation DAVIS 2017 (val) URVOS J&F 1st frame 44.1 #15 of 18 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) URVOS F 50.8 #32 of 33 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) URVOS J 47.0 #32 of 33 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) URVOS J&F 48.9 #32 of 33 Archive leaderboard report
Referring Video Object Segmentation MeViS URVOS F 29.9 #16 of 16 Archive leaderboard report
Referring Video Object Segmentation MeViS URVOS J 25.7 #16 of 16 Archive leaderboard report
Referring Video Object Segmentation MeViS URVOS J&F 27.8 #16 of 16 Archive leaderboard report
Referring Video Object Segmentation Ref-DAVIS17 URVOS F 56.0 #11 of 11 Archive leaderboard report
Referring Video Object Segmentation Ref-DAVIS17 URVOS J 47.3 #11 of 11 Archive leaderboard report
Referring Video Object Segmentation Ref-DAVIS17 URVOS J&F 51.6 #11 of 11 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections