{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/recurrent-multimodal-interaction-for","title":"Recurrent Multimodal Interaction for Referring Image Segmentation","arxiv_id":"1703.07939","date":"2017-03-23","proceeding":"ICCV 2017 10","authors":["Chenxi Liu","Zhe Lin","Xiaohui Shen","Jimei Yang","Xin Lu","Alan Yuille"],"abstract":"In this paper we are interested in the problem of image segmentation given\nnatural language descriptions, i.e. referring expressions. Existing works\ntackle this problem by first modeling images and sentences independently and\nthen segment images by combining these two types of representations. We argue\nthat learning word-to-image interaction is more native in the sense of jointly\nmodeling two modalities for the image segmentation task, and we propose\nconvolutional multimodal LSTM to encode the sequential interactions between\nindividual words, visual information, and spatial information. We show that our\nproposed model outperforms the baseline model on benchmark datasets. In\naddition, we analyze the intermediate output of the proposed multimodal LSTM\napproach and empirically explain how this approach enforces a more effective\nword-to-image interaction.","url_abs":"http://arxiv.org/abs/1703.07939v2","url_pdf":"http://arxiv.org/pdf/1703.07939v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"recurrent-multimodal-interaction-for","repo_url":"https://github.com/chenxi116/TF-phrasecut-public","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"image-segmentation","task_name":"Image Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"multimodal-interaction","task_name":"multimodal interaction"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1703.07939","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}