Papers › MaIL: A Unified Mask-Image-Language Trimodal Network for Referring Image Segmentation

MaIL: A Unified Mask-Image-Language Trimodal Network for Referring Image Segmentation

21 Nov 2021arXiv:2111.10747archive 2025-07-28

Zizhang Li, Mengmeng Wang, Jianbiao Mei, Yong liu

Referring image segmentation is a typical multi-modal task, which aims at generating a binary mask for referent described in given language expressions. Prior arts adopt a bimodal solution, taking images and languages as two modalities within an encoder-fusion-decoder pipeline. However, this pipeline is sub-optimal for the target task for two reasons. First, they only fuse high-level features produced by uni-modal encoders separately, which hinders sufficient cross-modal learning. Second, the uni-modal encoders are pre-trained independently, which brings inconsistency between pre-trained uni-modal tasks and the target multi-modal task. Besides, this pipeline often ignores or makes little use of intuitively beneficial instance-level features. To relieve these problems, we propose MaIL, which is a more concise encoder-decoder pipeline with a Mask-Image-Language trimodal encoder. Specifically, MaIL unifies uni-modal feature extractors and their fusion model into a deep modality interaction encoder, facilitating sufficient feature interaction across different modalities. Meanwhile, MaIL directly avoids the second limitation since no uni-modal encoders are needed anymore. Moreover, for the first time, we propose to introduce instance masks as an additional modality, which explicitly intensifies instance-level features and promotes finer segmentation results. The proposed MaIL set a new state-of-the-art on all frequently-used referring image segmentation datasets, including RefCOCO, RefCOCO+, and G-Ref, with significant gains, 3%-10% against previous best methods. Code will be released soon.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderImage SegmentationReferring Expression SegmentationSegmentationSemantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation G-Ref test A MaIL Overall IoU 62.87 #1 of 1 Archive leaderboard report
Referring Expression Segmentation G-Ref test B MaIL Overall IoU 61.81 #1 of 1 Archive leaderboard report
Referring Expression Segmentation G-Ref val MaIL Overall IoU 62.45 #1 of 1 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ test B MaIL Overall IoU 56.06 #18 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ testA MaIL Overall IoU 65.92 #21 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ val MaIL Overall IoU 62.23 #23 of 33 Archive leaderboard report
Referring Expression Segmentation RefCoCo val MaIL Overall IoU 70.13 #26 of 37 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections