{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/differential-attention-for-visual-question","title":"Differential Attention for Visual Question Answering","arxiv_id":"1804.00298","date":"2018-04-01","proceeding":"CVPR 2018 6","authors":["Badri Patro","Vinay P. Namboodiri"],"abstract":"In this paper we aim to answer questions based on images when provided with a\ndataset of question-answer pairs for a number of images during training. A\nnumber of methods have focused on solving this problem by using image based\nattention. This is done by focusing on a specific part of the image while\nanswering the question. Humans also do so when solving this problem. However,\nthe regions that the previous systems focus on are not correlated with the\nregions that humans focus on. The accuracy is limited due to this drawback. In\nthis paper, we propose to solve this problem by using an exemplar based method.\nWe obtain one or more supporting and opposing exemplars to obtain a\ndifferential attention region. This differential attention is closer to human\nattention than other image based attention methods. It also helps in obtaining\nimproved accuracy when answering questions. The method is evaluated on\nchallenging benchmark datasets. We perform better than other image based\nattention methods and are competitive with other state of the art methods that\nfocus on both image and questions.","url_abs":"http://arxiv.org/abs/1804.00298v2","url_pdf":"http://arxiv.org/pdf/1804.00298v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"differential-attention-for-visual-question","repo_url":"https://github.com/chirag26495/DAN_VQA","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.00298","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}