Papers › SynthRef: Generation of Synthetic Referring Expressions for Object Segmentation

SynthRef: Generation of Synthetic Referring Expressions for Object Segmentation

8 Jun 2021arXiv:2106.04403archive 2025-07-28

Ioannis Kazakos, Carles Ventura, Miriam Bellver, Carina Silberer, Xavier Giro-i-Nieto

Recent advances in deep learning have brought significant progress in visual grounding tasks such as language-guided video object segmentation. However, collecting large datasets for these tasks is expensive in terms of annotation time, which represents a bottleneck. To this end, we propose a novel method, namely SynthRef, for generating synthetic referring expressions for target objects in an image (or video frame), and we also present and disseminate the first large-scale dataset with synthetic referring expressions for video object segmentation. Our experiments demonstrate that by training with our synthetic referring expressions one can improve the ability of a model to generalize across different datasets, without any additional annotation cost. Moreover, our formulation allows its application to any object detection or segmentation dataset.

PaperPDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

imatge-upc/synthref officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ObjectReferring Expression SegmentationSegmentationVideo Object Segmentationobject-detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation DAVIS 2017 (val) RefVOS + SynthRef-YouTube-VIS J&F 1st frame 45.3 #12 of 18 Archive leaderboard report
Referring Expression Segmentation DAVIS 2017 (val) RefVOS + SynthRef-YouTube-VIS J&F Full video 44.8 #12 of 18 Archive leaderboard report
Referring Expression Segmentation DAVIS 2017 (val) RefVOS J&F 1st frame 45.1 #13 of 18 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS RefVOS-Human REs Mean IoU 39.5 #1 of 2 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS RefVOS-Human REs Precision@0.5 38.6 #1 of 2 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS RefVOS-Human REs Precision@0.9 6.9 #1 of 2 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS RefVOS-Synthetic REs Mean IoU 35.0 #2 of 2 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS RefVOS-Synthetic REs Precision@0.5 32.3 #2 of 2 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS RefVOS-Synthetic REs Precision@0.9 1.8 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections