{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dynamic-multimodal-instance-segmentation","title":"Dynamic Multimodal Instance Segmentation guided by natural language queries","arxiv_id":"1807.02257","date":"2018-07-06","proceeding":"ECCV 2018 9","authors":["Edgar Margffoy-Tuay","Juan C. Pérez","Emilio Botero","Pablo Arbeláez"],"abstract":"We address the problem of segmenting an object given a natural language\nexpression that describes it. Current techniques tackle this task by either\n(\\textit{i}) directly or recursively merging linguistic and visual information\nin the channel dimension and then performing convolutions; or by (\\textit{ii})\nmapping the expression to a space in which it can be thought of as a filter,\nwhose response is directly related to the presence of the object at a given\nspatial coordinate in the image, so that a convolution can be applied to look\nfor the object. We propose a novel method that integrates these two insights in\norder to fully exploit the recursive nature of language. Additionally, during\nthe upsampling process, we take advantage of the intermediate information\ngenerated when downsampling the image, so that detailed segmentations can be\nobtained. We compare our method against the state-of-the-art approaches in four\nstandard datasets, in which it surpasses all previous methods in six of eight\nof the splits for this task.","url_abs":"http://arxiv.org/abs/1807.02257v2","url_pdf":"http://arxiv.org/pdf/1807.02257v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dynamic-multimodal-instance-segmentation","repo_url":"https://github.com/BCV-Uniandes/query-objseg","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"dynamic-multimodal-instance-segmentation","repo_url":"https://github.com/andfoy/query-objseg","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"natural-language-queries","task_name":"Natural Language Queries"},{"task_slug":"object","task_name":"Object"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1807.02257","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}