{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-coreference-resolution-in-visual","title":"Visual Coreference Resolution in Visual Dialog using Neural Module Networks","arxiv_id":"1809.01816","date":"2018-09-06","proceeding":"ECCV 2018 9","authors":["Satwik Kottur","José M. F. Moura","Devi Parikh","Dhruv Batra","Marcus Rohrbach"],"abstract":"Visual dialog entails answering a series of questions grounded in an image,\nusing dialog history as context. In addition to the challenges found in visual\nquestion answering (VQA), which can be seen as one-round dialog, visual dialog\nencompasses several more. We focus on one such problem called visual\ncoreference resolution that involves determining which words, typically noun\nphrases and pronouns, co-refer to the same entity/object instance in an image.\nThis is crucial, especially for pronouns (e.g., `it'), as the dialog agent must\nfirst link it to a previous coreference (e.g., `boat'), and only then can rely\non the visual grounding of the coreference `boat' to reason about the pronoun\n`it'. Prior work (in visual dialog) models visual coreference resolution either\n(a) implicitly via a memory network over history, or (b) at a coarse level for\nthe entire question; and not explicitly at a phrase level of granularity. In\nthis work, we propose a neural module network architecture for visual dialog by\nintroducing two novel modules - Refer and Exclude - that perform explicit,\ngrounded, coreference resolution at a finer word level. We demonstrate the\neffectiveness of our model on MNIST Dialog, a visually simple yet\ncoreference-wise complex dataset, by achieving near perfect accuracy, and on\nVisDial, a large and challenging visual dialog dataset on real images, where\nour model outperforms other approaches, and is more interpretable, grounded,\nand consistent qualitatively.","url_abs":"http://arxiv.org/abs/1809.01816v1","url_pdf":"http://arxiv.org/pdf/1809.01816v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visual-coreference-resolution-in-visual","repo_url":"https://github.com/facebookresearch/corefnmn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"common-sense-reasoning","task_name":"Common Sense Reasoning"},{"task_slug":"coreference-resolution","task_name":"Coreference Resolution"},{"task_slug":"visual-dialogue","task_name":"Visual Dialog"},{"task_slug":"visual-grounding","task_name":"Visual Grounding"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"},{"task_slug":"coreference-resolution-1","task_name":"coreference-resolution"}],"methods":[{"method_slug":"memory-network","method_name":"Memory Network"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/common-sense-reasoning-on-visual-dialog-v0-9","task":"Common Sense Reasoning","dataset":"Visual Dialog v0.9","model":"NMN [kottur2018visual]","rank_in_archive_order":1,"of":1,"metrics":{"1 in 10 R@5":"80.1"},"uses_additional_data":false},{"leaderboard":"/sota/visual-dialog-on-visdial-v09-val","task":"Visual Dialog","dataset":"VisDial v0.9 val","model":"CorefNMN (ResNet-152)","rank_in_archive_order":3,"of":18,"metrics":{"MRR":"64.1","Mean Rank":"4.45","R@1":"50.92","R@10":"88.81","R@5":"80.18"},"uses_additional_data":false},{"leaderboard":"/sota/visual-dialog-on-visdial-v09-val","task":"Visual Dialog","dataset":"VisDial v0.9 val","model":"CorefNMN","rank_in_archive_order":5,"of":18,"metrics":{"MRR":"63.6","Mean Rank":"4.53","R@1":"50.24","R@10":"88.51","R@5":"79.81"},"uses_additional_data":false},{"leaderboard":"/sota/visual-dialog-on-visual-dialog-v1-0-test-std","task":"Visual Dialog","dataset":"Visual Dialog v1.0 test-std","model":"CorefNMN (ResNet-152)","rank_in_archive_order":67,"of":80,"metrics":{"MRR (x 100)":"61.50","Mean":"4.40","NDCG (x 100)":"54.70","R@1":"47.55","R@10":"88.80","R@5":"78.10"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1809.01816","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}