Papers › RefVOS: A Closer Look at Referring Expressions for Video Object Segmentation

RefVOS: A Closer Look at Referring Expressions for Video Object Segmentation

1 Oct 2020arXiv:2010.00263archive 2025-07-28

Miriam Bellver, Carles Ventura, Carina Silberer, Ioannis Kazakos, Jordi Torres, Xavier Giro-i-Nieto

The task of video object segmentation with referring expressions (language-guided VOS) is to, given a linguistic phrase and a video, generate binary masks for the object to which the phrase refers. Our work argues that existing benchmarks used for this task are mainly composed of trivial cases, in which referents can be identified with simple phrases. Our analysis relies on a new categorization of the phrases in the DAVIS-2017 and Actor-Action datasets into trivial and non-trivial REs, with the non-trivial REs annotated with seven RE semantic categories. We leverage this data to analyze the results of RefVOS, a novel neural network that obtains competitive results for the task of language-guided image segmentation and state of the art results for language-guided VOS. Our study indicates that the major challenges for the task are related to understanding motion and static actions.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

miriambellver/refvos officialmentioned in papermentioned on GitHubpytorch report
imatge-upc/refvos mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image SegmentationReferring Expression SegmentationSegmentationVideo Object Segmentation

Datasets

Introduced by this paper, per the archive.

A2DreA2Dre+

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation A2D Sentences RefVOS IoU mean 0.599 #27 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences RefVOS IoU overall 0.599 #27 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences RefVOS Precision@0.5 0.495 #27 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences RefVOS Precision@0.9 0.064 #27 of 27 Archive leaderboard report
Referring Expression Segmentation A2Dre test RefVos Mean IoU 33.2 #1 of 1 Archive leaderboard report
Referring Expression Segmentation A2Dre test RefVos Overall IoU 47.5 #1 of 1 Archive leaderboard report
Referring Expression Segmentation DAVIS 2017 (val) RefVOS J&F 1st frame 44.5 #14 of 18 Archive leaderboard report
Referring Expression Segmentation DAVIS 2017 (val) RefVOS J&F Full video 45.1 #14 of 18 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ test B RefVOS with BERT + MLM loss Overall IoU 36.17 #28 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ testA RefVOS with BERT + MLM Loss Overall IoU 49.73 #28 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ val RefVOS with BERT + MLM loss Overall IoU 44.71 #31 of 33 Archive leaderboard report
Referring Expression Segmentation RefCoCo val RefVOS with BERT + MLM loss Overall IoU 59.45 #31 of 37 Archive leaderboard report
Referring Expression Segmentation RefCoCo val RefVOS with BERT Pre-train Overall IoU 58.65 #33 of 37 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

1x1 ConvolutionAdamAttentionAttention DropoutBERTConvolutionDense ConnectionsDilated ConvolutionDropoutGrouped ConvolutionLayer NormalizationLinear LayerLinear Warmup With Linear DecayMulti-Head AttentionMultiscale Dilated Convolution BlockResidual ConnectionSoftmaxVOSWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections