Papers › MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation

MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation

28 Nov 2024arXiv:2411.19067archive 2025-07-28

Minhyun Lee, Seungho Lee, Song Park, Dongyoon Han, Byeongho Heo, Hyunjung Shim

Referring Image Segmentation (RIS) is an advanced vision-language task that involves identifying and segmenting objects within an image as described by free-form text descriptions. While previous studies focused on aligning visual and language features, exploring training techniques, such as data augmentation, remains underexplored. In this work, we explore effective data augmentation for RIS and propose a novel training framework called Masked Referring Image Segmentation (MaskRIS). We observe that the conventional image augmentations fall short of RIS, leading to performance degradation, while simple random masking significantly enhances the performance of RIS. MaskRIS uses both image and text masking, followed by Distortion-aware Contextual Learning (DCL) to fully exploit the benefits of the masking strategy. This approach can improve the model's robustness to occlusions, incomplete information, and various linguistic complexities, resulting in a significant performance improvement. Experiments demonstrate that MaskRIS can easily be applied to various RIS models, outperforming existing methods in both fully supervised and weakly supervised settings. Finally, MaskRIS achieves new state-of-the-art performance on RefCOCO, RefCOCO+, and RefCOCOg datasets. Code is available at https://github.com/naver-ai/maskris.

PaperPDFCode

Code

naver-ai/maskris officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Data AugmentationImage SegmentationReferring Expression Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation RefCOCO testA MaskRIS (Swin-B, combined DB) Overall IoU 80.64 #7 of 13 Archive leaderboard report
Referring Expression Segmentation RefCOCO testA MaskRIS (Swin-B) Mean IoU 80.24 #8 of 13 Archive leaderboard report
Referring Expression Segmentation RefCOCO testA MaskRIS (Swin-B) Overall IoU 78.96 #8 of 13 Archive leaderboard report
Referring Expression Segmentation RefCOCO testB MaskRIS (Swin-B, combined DB) Overall IoU 75.1 #7 of 13 Archive leaderboard report
Referring Expression Segmentation RefCOCO testB MaskRIS (Swin-B) Mean IoU 76.06 #8 of 13 Archive leaderboard report
Referring Expression Segmentation RefCOCO testB MaskRIS (Swin-B) Overall IoU 73.96 #8 of 13 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ test B MaskRIS (Swin-B, combined DB) Overall IoU 62.83 #11 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ test B MaskRIS (Swin-B) Mean IoU 64.5 #13 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ test B MaskRIS (Swin-B) Overall IoU 59.39 #13 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ testA MaskRIS (Swin-B, combined DB) Overall IoU 75.15 #10 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ testA MaskRIS (Swin-B) Mean IoU 76.73 #14 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ testA MaskRIS (Swin-B) Overall IoU 74.46 #14 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ val MaskRIS (Swin-B, combined DB) Overall IoU 70.26 #14 of 33 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ val MaskRIS (Swin-B) Mean IoU 71.68 #18 of 33 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ val MaskRIS (Swin-B) Overall IoU 67.54 #18 of 33 Archive leaderboard report
Referring Expression Segmentation RefCOCOg-test MaskRIS (Swin-B, combined DB) Overall IoU 71.09 #9 of 18 Archive leaderboard report
Referring Expression Segmentation RefCOCOg-test MaskRIS (Swin-B) Mean IoU 69.42 #13 of 18 Archive leaderboard report
Referring Expression Segmentation RefCOCOg-test MaskRIS (Swin-B) Overall IoU 66.5 #13 of 18 Archive leaderboard report
Referring Expression Segmentation RefCOCOg-val MaskRIS (Swin-B, combined DB) Overall IoU 69.12 #13 of 23 Archive leaderboard report
Referring Expression Segmentation RefCOCOg-val MaskRIS (Swin-B) Mean IoU 69.31 #15 of 23 Archive leaderboard report
Referring Expression Segmentation RefCOCOg-val MaskRIS (Swin-B) Overall IoU 65.55 #15 of 23 Archive leaderboard report
Referring Expression Segmentation RefCoCo val MaskRIS (Swin-B, combined DB) Overall IoU 78.71 #13 of 37 Archive leaderboard report
Referring Expression Segmentation RefCoCo val MaskRIS (Swin-B) Mean IoU 78.35 #16 of 37 Archive leaderboard report
Referring Expression Segmentation RefCoCo val MaskRIS (Swin-B) Overall IoU 76.49 #16 of 37 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionAttention DropoutBERTBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerLinear Warmup With Linear DecayMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxStochastic DepthSwin TransformerTransformerVision TransformerWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections