{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/grounding-referring-expressions-in-images-by","title":"Grounding Referring Expressions in Images by Variational Context","arxiv_id":"1712.01892","date":"2017-12-05","proceeding":"CVPR 2018 6","authors":["Hanwang Zhang","Yulei Niu","Shih-Fu Chang"],"abstract":"We focus on grounding (i.e., localizing or linking) referring expressions in\nimages, e.g., \"largest elephant standing behind baby elephant\". This is a\ngeneral yet challenging vision-language task since it does not only require the\nlocalization of objects, but also the multimodal comprehension of context ---\nvisual attributes (e.g., \"largest\", \"baby\") and relationships (e.g., \"behind\")\nthat help to distinguish the referent from other objects, especially those of\nthe same category. Due to the exponential complexity involved in modeling the\ncontext associated with multiple image regions, existing work oversimplifies\nthis task to pairwise region modeling by multiple instance learning. In this\npaper, we propose a variational Bayesian method, called Variational Context, to\nsolve the problem of complex context modeling in referring expression\ngrounding. Our model exploits the reciprocal relation between the referent and\ncontext, i.e., either of them influences the estimation of the posterior\ndistribution of the other, and thereby the search space of context can be\ngreatly reduced, resulting in better localization of referent. We develop a\nnovel cue-specific language-vision embedding network that learns this\nreciprocity model end-to-end. We also extend the model to the unsupervised\nsetting where no annotation for the referent is available. Extensive\nexperiments on various benchmarks show consistent improvement over\nstate-of-the-art methods in both supervised and unsupervised settings.","url_abs":"http://arxiv.org/abs/1712.01892v2","url_pdf":"http://arxiv.org/pdf/1712.01892v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"grounding-referring-expressions-in-images-by","repo_url":"https://github.com/yuleiniu/vc","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"multiple-instance-learning","task_name":"Multiple Instance Learning"},{"task_slug":"referring-expression","task_name":"Referring Expression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1712.01892","atlas_url":"https://app.syntology.ai/?focus=1712.01892","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}