{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/modeling-relationships-in-referential","title":"Modeling Relationships in Referential Expressions with Compositional Modular Networks","arxiv_id":"1611.09978","date":"2016-11-30","proceeding":"CVPR 2017 7","authors":["Ronghang Hu","Marcus Rohrbach","Jacob Andreas","Trevor Darrell","Kate Saenko"],"abstract":"People often refer to entities in an image in terms of their relationships\nwith other entities. For example, \"the black cat sitting under the table\"\nrefers to both a \"black cat\" entity and its relationship with another \"table\"\nentity. Understanding these relationships is essential for interpreting and\ngrounding such natural language expressions. Most prior work focuses on either\ngrounding entire referential expressions holistically to one region, or\nlocalizing relationships based on a fixed set of categories. In this paper we\ninstead present a modular deep architecture capable of analyzing referential\nexpressions into their component parts, identifying entities and relationships\nmentioned in the input expression and grounding them all in the scene. We call\nthis approach Compositional Modular Networks (CMNs): a novel architecture that\nlearns linguistic analysis and visual inference end-to-end. Our approach is\nbuilt around two types of neural modules that inspect local regions and\npairwise interactions between regions. We evaluate CMNs on multiple referential\nexpression datasets, outperforming state-of-the-art approaches on all tasks.","url_abs":"http://arxiv.org/abs/1611.09978v1","url_pdf":"http://arxiv.org/pdf/1611.09978v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"modeling-relationships-in-referential","repo_url":"https://github.com/hengyuan-hu/bottom-up-attention-vqa","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"modeling-relationships-in-referential","repo_url":"https://github.com/thilinicooray/Bottom-up-vqa","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-question-answering-on-visual-genome-1","task":"Visual Question Answering (VQA)","dataset":"Visual Genome (pairs)","model":"CMN","rank_in_archive_order":1,"of":1,"metrics":{"Percentage correct":"28.52"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-visual-genome","task":"Visual Question Answering (VQA)","dataset":"Visual Genome (subjects)","model":"CMN","rank_in_archive_order":1,"of":1,"metrics":{"Percentage correct":"44.24"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-visual7w","task":"Visual Question Answering (VQA)","dataset":"Visual7W","model":"CMN","rank_in_archive_order":1,"of":4,"metrics":{"Percentage correct":"72.53"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1611.09978","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}