{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/segmentation-from-natural-language","title":"Segmentation from Natural Language Expressions","arxiv_id":"1603.06180","date":"2016-03-20","proceeding":null,"authors":["Ronghang Hu","Marcus Rohrbach","Trevor Darrell"],"abstract":"In this paper we approach the novel problem of segmenting an image based on a\nnatural language expression. This is different from traditional semantic\nsegmentation over a predefined set of semantic classes, as e.g., the phrase\n\"two men sitting on the right bench\" requires segmenting only the two people on\nthe right bench and no one standing or sitting on another bench. Previous\napproaches suitable for this task were limited to a fixed set of categories\nand/or rectangular regions. To produce pixelwise segmentation for the language\nexpression, we propose an end-to-end trainable recurrent and convolutional\nnetwork model that jointly learns to process visual and linguistic information.\nIn our model, a recurrent LSTM network is used to encode the referential\nexpression into a vector representation, and a fully convolutional network is\nused to a extract a spatial feature map from the image and output a spatial\nresponse map for the target object. We demonstrate on a benchmark dataset that\nour model can produce quality segmentation output from the natural language\nexpression, and outperforms baseline methods by a large margin.","url_abs":"http://arxiv.org/abs/1603.06180v1","url_pdf":"http://arxiv.org/pdf/1603.06180v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"segmentation-from-natural-language","repo_url":"https://github.com/ronghanghu/text_objseg","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"segmentation-from-natural-language","repo_url":"https://github.com/ssharpe42/NLQAC_Instance_Selection","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"segmentation-from-natural-language","repo_url":"https://github.com/ssharpe42/NLQAC_ObjSeg","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"segmentation-from-natural-language","repo_url":"https://github.com/ssharpe42/VNLQAC","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"referring-expression-segmentation","task_name":"Referring Expression Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/referring-expression-segmentation-on-a2d","task":"Referring Expression Segmentation","dataset":"A2D Sentences","model":"Hu et al.","rank_in_archive_order":22,"of":27,"metrics":{"AP":"0.132","IoU mean":"0.350","IoU overall":"0.474","Precision@0.5":"0.348","Precision@0.6":"0.236","Precision@0.7":"0.133","Precision@0.8":"0.033","Precision@0.9":"0.000"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-j-hmdb","task":"Referring Expression Segmentation","dataset":"J-HMDB","model":"Hu et al.","rank_in_archive_order":16,"of":21,"metrics":{"AP":"0.178","IoU mean":"0.528","IoU overall":"0.546","Precision@0.5":"0.633","Precision@0.6":"0.350","Precision@0.7":"0.085","Precision@0.8":"0.002","Precision@0.9":"0.000"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1603.06180","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}