Papers › Cross-Modal Self-Attention Network for Referring Image Segmentation

Cross-Modal Self-Attention Network for Referring Image Segmentation

9 Apr 2019CVPR 2019 6arXiv:1904.04745archive 2025-07-28

Linwei Ye, Mrigank Rochan, Zhi Liu, Yang Wang

We consider the problem of referring image segmentation. Given an input image and a natural language expression, the goal is to segment the object referred by the language expression in the image. Existing works in this area treat the language expression and the input image separately in their representations. They do not sufficiently capture long-range correlations between these two modalities. In this paper, we propose a cross-modal self-attention (CMSA) module that effectively captures the long-range dependencies between linguistic and visual features. Our model can adaptively focus on informative words in the referring expression and important regions in the input image. In addition, we propose a gated multi-level fusion module to selectively integrate self-attentive cross-modal features corresponding to different levels in the image. This module controls the information flow of features at different levels. We validate the proposed approach on four evaluation datasets. Our proposed approach consistently outperforms existing state-of-the-art methods.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

lwye/CMSA-Net mentioned on GitHubtf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image SegmentationReferring ExpressionReferring Expression SegmentationReferring Video Object SegmentationSemantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation RefCOCO+ test B CMSA Overall IoU 37.89 #27 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ testA CMSA Overall IoU 47.60 #29 of 30 Archive leaderboard report
Referring Expression Segmentation RefCOCO+ val CMSA Overall IoU 43.76 #32 of 33 Archive leaderboard report
Referring Expression Segmentation RefCoCo val CMSA Overall IoU 58.32 #34 of 37 Archive leaderboard report
Referring Video Object Segmentation Refer-YouTube-VOS CMSA F 38.1 #18 of 18 Archive leaderboard report
Referring Video Object Segmentation Refer-YouTube-VOS CMSA J 34.8 #18 of 18 Archive leaderboard report
Referring Video Object Segmentation Refer-YouTube-VOS CMSA J&F 36.4 #18 of 18 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections