Papers › Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding...

Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding

12 Apr 2024arXiv:2404.08590archive 2025-07-28

Hai Nguyen-Truong, E-Ro Nguyen, Tuan-Anh Vu, Minh-Triet Tran, Binh-Son Hua, Sai-Kit Yeung

Referring image segmentation is a challenging task that involves generating pixel-wise segmentation masks based on natural language descriptions. The complexity of this task increases with the intricacy of the sentences provided. Existing methods have relied mostly on visual features to generate the segmentation masks while treating text features as supporting components. However, this under-utilization of text understanding limits the model's capability to fully comprehend the given expressions. In this work, we propose a novel framework that specifically emphasizes object and context comprehension inspired by human cognitive processes through Vision-Aware Text Features. Firstly, we introduce a CLIP Prior module to localize the main object of interest and embed the object heatmap into the query initialization process. Secondly, we propose a combination of two components: Contextual Multimodal Decoder and Meaning Consistency Constraint, to further enhance the coherent and consistent interpretation of language cues with the contextual understanding obtained from the image. Our method achieves significant performance improvements on three benchmark datasets RefCOCO, RefCOCO+ and G-Ref. Project page: \url{https://vatex.hkustvgd.com/}.

PaperPDFCode

Code

nero1342/VATEX mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderImage SegmentationObjectReferring Expression SegmentationReferring Video Object SegmentationSegmentationSemantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation DAVIS 2017 (val) VATEX J&F score 65.4 #18 of 18 Archive leaderboard report
Referring Expression Segmentation RefCOCO testA VATEX mIoU 79.64 #13 of 13 Archive leaderboard report
Referring Expression Segmentation RefCOCO testB VATEX mIoU 75.64 #13 of 13 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ test B VATEX mIoU 62.52 #30 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ testA VATEX mIoU 74.41 #30 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ val VATEX Mean IoU 70.02 #33 of 33 Archive leaderboard report
Referring Expression Segmentation RefCOCOg-test VATEX mIoU 70.58 #18 of 18 Archive leaderboard report
Referring Expression Segmentation RefCOCOg-val VATEX IoU 0.7554 #23 of 23 Archive leaderboard report
Referring Expression Segmentation RefCOCOg-val VATEX mIoU 69.73 #23 of 23 Archive leaderboard report
Referring Expression Segmentation RefCoCo val VATEX mIoU 78.16 #37 of 37 Archive leaderboard report
Referring Video Object Segmentation Refer-YouTube-VOS VATEX F 67.5 #8 of 18 Archive leaderboard report
Referring Video Object Segmentation Refer-YouTube-VOS VATEX J 63.3 #8 of 18 Archive leaderboard report
Referring Video Object Segmentation Refer-YouTube-VOS VATEX J&F 65.4 #8 of 18 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CLIPHeatmap

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections