Papers › SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation

SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation

2 Jul 2024arXiv:2407.02389archive 2025-07-28

Sayan Nag, Koustava Goswami, Srikrishna Karanam

Referring Expression Segmentation (RES) aims to provide a segmentation mask of the target object in an image referred to by the text (i.e., referring expression). Existing methods require large-scale mask annotations. Moreover, such approaches do not generalize well to unseen/zero-shot scenarios. To address the aforementioned issues, we propose a weakly-supervised bootstrapping architecture for RES with several new algorithmic innovations. To the best of our knowledge, ours is the first approach that considers only a fraction of both mask and box annotations (shown in Figure 1 and Table 1) for training. To enable principled training of models in such low-annotation settings, improve image-text region-level alignment, and further enhance spatial localization of the target object in the image, we propose Cross-modal Fusion with Attention Consistency module. For automatic pseudo-labeling of unlabeled samples, we introduce a novel Mask Validity Filtering routine based on a spatially aware zero-shot proposal scoring approach. Extensive experiments show that with just 30% annotations, our model SafaRi achieves 59.31 and 48.26 mIoUs as compared to 58.93 and 48.19 mIoUs obtained by the fully-supervised SOTA method SeqTR respectively on RefCOCO+@testA and RefCOCO+testB datasets. SafaRi also outperforms SeqTR by 11.7% (on RefCOCO+testA) and 19.6% (on RefCOCO+testB) in a fully-supervised setting and demonstrates strong generalization capabilities in unseen/zero-shot tasks.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Referring ExpressionReferring Expression SegmentationWeakly Supervised Referring Expression Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation DAVIS 2017 (val) SafaRi-B J&F 1st frame 61.3 #6 of 18 Archive leaderboard report
Referring Expression Segmentation DAVIS 2017 (val) SafaRi-B Zero-Shot Transfer true #6 of 18 Archive leaderboard report
Referring Expression Segmentation RefCOCO testA SafaRi Overall IoU 77.83 #11 of 13 Archive leaderboard report
Referring Expression Segmentation RefCOCO testB SafaRi Overall IoU 70.71 #11 of 13 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ test B SafaRi-B Overall IoU 64.88 #10 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ testA SafaRi-B Overall IoU 74.53 #13 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ val SafaRi-B Overall IoU 70.78 #12 of 33 Archive leaderboard report
Referring Expression Segmentation RefCOCOg-test SafaRi-B Overall IoU 71.06 #10 of 18 Archive leaderboard report
Referring Expression Segmentation RefCOCOg-val SafaRi-B Overall IoU 70.48 #11 of 23 Archive leaderboard report
Referring Expression Segmentation RefCoCo val SafaRi-B Overall IoU 77.21 #15 of 37 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AWAREAttentionSoftmax

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections