{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-reference-resolution-using-attention","title":"Visual Reference Resolution using Attention Memory for Visual Dialog","arxiv_id":"1709.07992","date":"2017-09-23","proceeding":"NeurIPS 2017 12","authors":["Paul Hongsuck Seo","Andreas Lehrmann","Bohyung Han","Leonid Sigal"],"abstract":"Visual dialog is a task of answering a series of inter-dependent questions\ngiven an input image, and often requires to resolve visual references among the\nquestions. This problem is different from visual question answering (VQA),\nwhich relies on spatial attention (a.k.a. visual grounding) estimated from an\nimage and question pair. We propose a novel attention mechanism that exploits\nvisual attentions in the past to resolve the current reference in the visual\ndialog scenario. The proposed model is equipped with an associative attention\nmemory storing a sequence of previous (attention, key) pairs. From this memory,\nthe model retrieves the previous attention, taking into account recency, which\nis most relevant for the current question, in order to resolve potentially\nambiguous references. The model then merges the retrieved attention with a\ntentative one to obtain the final attention for the current question;\nspecifically, we use dynamic parameter prediction to combine the two attentions\nconditioned on the question. Through extensive experiments on a new synthetic\nvisual dialog dataset, we show that our model significantly outperforms the\nstate-of-the-art (by ~16 % points) in situations, where visual reference\nresolution plays an important role. Moreover, the proposed model achieves\nsuperior performance (~ 2 % points improvement) in the Visual Dialog dataset,\ndespite having significantly fewer parameters than the baselines.","url_abs":"http://arxiv.org/abs/1709.07992v3","url_pdf":"http://arxiv.org/pdf/1709.07992v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"parameter-prediction","task_name":"Parameter Prediction"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-dialogue","task_name":"Visual Dialog"},{"task_slug":"visual-grounding","task_name":"Visual Grounding"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-dialog-on-visdial-v09-val","task":"Visual Dialog","dataset":"VisDial v0.9 val","model":"AMEM","rank_in_archive_order":18,"of":18,"metrics":{"Mean Rank":"4.86","R@1":"48.53","R@10":"87.43","R@5":"78.66"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1709.07992","atlas_url":"https://app.syntology.ai/?focus=1709.07992","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}