{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/safari-adaptive-sequence-transformer-for","title":"SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation","arxiv_id":"2407.02389","date":"2024-07-02","proceeding":null,"authors":["Sayan Nag","Koustava Goswami","Srikrishna Karanam"],"abstract":"Referring Expression Segmentation (RES) aims to provide a segmentation mask of the target object in an image referred to by the text (i.e., referring expression). Existing methods require large-scale mask annotations. Moreover, such approaches do not generalize well to unseen/zero-shot scenarios. To address the aforementioned issues, we propose a weakly-supervised bootstrapping architecture for RES with several new algorithmic innovations. To the best of our knowledge, ours is the first approach that considers only a fraction of both mask and box annotations (shown in Figure 1 and Table 1) for training. To enable principled training of models in such low-annotation settings, improve image-text region-level alignment, and further enhance spatial localization of the target object in the image, we propose Cross-modal Fusion with Attention Consistency module. For automatic pseudo-labeling of unlabeled samples, we introduce a novel Mask Validity Filtering routine based on a spatially aware zero-shot proposal scoring approach. Extensive experiments show that with just 30% annotations, our model SafaRi achieves 59.31 and 48.26 mIoUs as compared to 58.93 and 48.19 mIoUs obtained by the fully-supervised SOTA method SeqTR respectively on RefCOCO+@testA and RefCOCO+testB datasets. SafaRi also outperforms SeqTR by 11.7% (on RefCOCO+testA) and 19.6% (on RefCOCO+testB) in a fully-supervised setting and demonstrates strong generalization capabilities in unseen/zero-shot tasks.","url_abs":"https://arxiv.org/abs/2407.02389v1","url_pdf":"https://arxiv.org/pdf/2407.02389v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"referring-expression","task_name":"Referring Expression"},{"task_slug":"referring-expression-segmentation","task_name":"Referring Expression Segmentation"},{"task_slug":"weakly-supervised-referring-expression","task_name":"Weakly Supervised Referring Expression Segmentation"}],"methods":[{"method_slug":"aware","method_name":"AWARE"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/referring-expression-segmentation-on-davis","task":"Referring Expression Segmentation","dataset":"DAVIS 2017 (val)","model":"SafaRi-B","rank_in_archive_order":6,"of":18,"metrics":{"J&F 1st frame":"61.3","Zero-Shot Transfer":"true"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-8","task":"Referring Expression Segmentation","dataset":"RefCOCO testA","model":"SafaRi","rank_in_archive_order":11,"of":13,"metrics":{"Overall IoU":"77.83"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-9","task":"Referring Expression Segmentation","dataset":"RefCOCO testB","model":"SafaRi","rank_in_archive_order":11,"of":13,"metrics":{"Overall IoU":"70.71"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-5","task":"Referring Expression Segmentation","dataset":"RefCOCO+ test B","model":"SafaRi-B","rank_in_archive_order":10,"of":30,"metrics":{"Overall IoU":"64.88"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-4","task":"Referring Expression Segmentation","dataset":"RefCOCO+ testA","model":"SafaRi-B","rank_in_archive_order":13,"of":30,"metrics":{"Overall IoU":"74.53"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-3","task":"Referring Expression Segmentation","dataset":"RefCOCO+ val","model":"SafaRi-B","rank_in_archive_order":12,"of":33,"metrics":{"Overall IoU":"70.78"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refcocog-1","task":"Referring Expression Segmentation","dataset":"RefCOCOg-test","model":"SafaRi-B","rank_in_archive_order":10,"of":18,"metrics":{"Overall IoU":"71.06"},"uses_additional_data":true},{"leaderboard":"/sota/referring-expression-segmentation-on-refcocog","task":"Referring Expression Segmentation","dataset":"RefCOCOg-val","model":"SafaRi-B","rank_in_archive_order":11,"of":23,"metrics":{"Overall IoU":"70.48"},"uses_additional_data":true},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco","task":"Referring Expression Segmentation","dataset":"RefCoCo val","model":"SafaRi-B","rank_in_archive_order":15,"of":37,"metrics":{"Overall IoU":"77.21"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2407.02389","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}