{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/recursive-visual-attention-in-visual-dialog","title":"Recursive Visual Attention in Visual Dialog","arxiv_id":"1812.02664","date":"2018-12-06","proceeding":"CVPR 2019 6","authors":["Yulei Niu","Hanwang Zhang","Manli Zhang","Jianhong Zhang","Zhiwu Lu","Ji-Rong Wen"],"abstract":"Visual dialog is a challenging vision-language task, which requires the agent\nto answer multi-round questions about an image. It typically needs to address\ntwo major problems: (1) How to answer visually-grounded questions, which is the\ncore challenge in visual question answering (VQA); (2) How to infer the\nco-reference between questions and the dialog history. An example of visual\nco-reference is: pronouns (\\eg, ``they'') in the question (\\eg, ``Are they on\nor off?'') are linked with nouns (\\eg, ``lamps'') appearing in the dialog\nhistory (\\eg, ``How many lamps are there?'') and the object grounded in the\nimage. In this work, to resolve the visual co-reference for visual dialog, we\npropose a novel attention mechanism called Recursive Visual Attention (RvA).\nSpecifically, our dialog agent browses the dialog history until the agent has\nsufficient confidence in the visual co-reference resolution, and refines the\nvisual attention recursively. The quantitative and qualitative experimental\nresults on the large-scale VisDial v0.9 and v1.0 datasets demonstrate that the\nproposed RvA not only outperforms the state-of-the-art methods, but also\nachieves reasonable recursion and interpretable attention maps without\nadditional annotations. The code is available at\n\\url{https://github.com/yuleiniu/rva}.","url_abs":"http://arxiv.org/abs/1812.02664v2","url_pdf":"http://arxiv.org/pdf/1812.02664v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"recursive-visual-attention-in-visual-dialog","repo_url":"https://github.com/yuleiniu/rva","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-dialogue","task_name":"Visual Dialog"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-dialog-on-visdial-v09-val","task":"Visual Dialog","dataset":"VisDial v0.9 val","model":"RVA","rank_in_archive_order":13,"of":18,"metrics":{"MRR":"0.6634","Mean Rank":"3.93","R@1":"52.71","R@10":"90.73","R@5":"82.97"},"uses_additional_data":false},{"leaderboard":"/sota/visual-dialog-on-visual-dialog-v1-0-test-std","task":"Visual Dialog","dataset":"Visual Dialog v1.0 test-std","model":"RVA","rank_in_archive_order":65,"of":80,"metrics":{"MRR (x 100)":"63.03","Mean":"4.18","NDCG (x 100)":"55.59","R@1":"49.03","R@10":"89.83","R@5":"80.40"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.02664","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}