{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cross-modal-self-attention-network-for","title":"Cross-Modal Self-Attention Network for Referring Image Segmentation","arxiv_id":"1904.04745","date":"2019-04-09","proceeding":"CVPR 2019 6","authors":["Linwei Ye","Mrigank Rochan","Zhi Liu","Yang Wang"],"abstract":"We consider the problem of referring image segmentation. Given an input image\nand a natural language expression, the goal is to segment the object referred\nby the language expression in the image. Existing works in this area treat the\nlanguage expression and the input image separately in their representations.\nThey do not sufficiently capture long-range correlations between these two\nmodalities. In this paper, we propose a cross-modal self-attention (CMSA)\nmodule that effectively captures the long-range dependencies between linguistic\nand visual features. Our model can adaptively focus on informative words in the\nreferring expression and important regions in the input image. In addition, we\npropose a gated multi-level fusion module to selectively integrate\nself-attentive cross-modal features corresponding to different levels in the\nimage. This module controls the information flow of features at different\nlevels. We validate the proposed approach on four evaluation datasets. Our\nproposed approach consistently outperforms existing state-of-the-art methods.","url_abs":"http://arxiv.org/abs/1904.04745v1","url_pdf":"http://arxiv.org/pdf/1904.04745v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cross-modal-self-attention-network-for","repo_url":"https://github.com/lwye/CMSA-Net","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-segmentation","task_name":"Image Segmentation"},{"task_slug":"referring-expression","task_name":"Referring Expression"},{"task_slug":"referring-expression-segmentation","task_name":"Referring Expression Segmentation"},{"task_slug":"referring-video-object-segmentation","task_name":"Referring Video Object Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-5","task":"Referring Expression Segmentation","dataset":"RefCOCO+ test B","model":"CMSA","rank_in_archive_order":27,"of":30,"metrics":{"Overall IoU":"37.89"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-4","task":"Referring Expression Segmentation","dataset":"RefCOCO+ testA","model":"CMSA","rank_in_archive_order":29,"of":30,"metrics":{"Overall IoU":"47.60"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-3","task":"Referring Expression Segmentation","dataset":"RefCOCO+ val","model":"CMSA","rank_in_archive_order":32,"of":33,"metrics":{"Overall IoU":"43.76"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco","task":"Referring Expression Segmentation","dataset":"RefCoCo val","model":"CMSA","rank_in_archive_order":34,"of":37,"metrics":{"Overall IoU":"58.32"},"uses_additional_data":false},{"leaderboard":"/sota/referring-video-object-segmentation-on-refer","task":"Referring Video Object Segmentation","dataset":"Refer-YouTube-VOS","model":"CMSA","rank_in_archive_order":18,"of":18,"metrics":{"F":"38.1","J":"34.8","J&F":"36.4"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.04745","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}