{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/making-the-v-in-vqa-matter-elevating-the-role","title":"Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering","arxiv_id":"1612.00837","date":"2016-12-02","proceeding":"CVPR 2017 7","authors":["Yash Goyal","Tejas Khot","Douglas Summers-Stay","Dhruv Batra","Devi Parikh"],"abstract":"Problems at the intersection of vision and language are of significant\nimportance both as challenging research questions and for the rich set of\napplications they enable. However, inherent structure in our world and bias in\nour language tend to be a simpler signal for learning than visual modalities,\nresulting in models that ignore visual information, leading to an inflated\nsense of their capability.\n  We propose to counter these language priors for the task of Visual Question\nAnswering (VQA) and make vision (the V in VQA) matter! Specifically, we balance\nthe popular VQA dataset by collecting complementary images such that every\nquestion in our balanced dataset is associated with not just a single image,\nbut rather a pair of similar images that result in two different answers to the\nquestion. Our dataset is by construction more balanced than the original VQA\ndataset and has approximately twice the number of image-question pairs. Our\ncomplete balanced dataset is publicly available at www.visualqa.org as part of\nthe 2nd iteration of the Visual Question Answering Dataset and Challenge (VQA\nv2.0).\n  We further benchmark a number of state-of-art VQA models on our balanced\ndataset. All models perform significantly worse on our balanced dataset,\nsuggesting that these models have indeed learned to exploit language priors.\nThis finding provides the first concrete empirical evidence for what seems to\nbe a qualitative sense among practitioners.\n  Finally, our data collection protocol for identifying complementary images\nenables us to develop a novel interpretable model, which in addition to\nproviding an answer to the given (image, question) pair, also provides a\ncounter-example based explanation. Specifically, it identifies an image that is\nsimilar to the original image, but it believes has a different answer to the\nsame question. This can help in building trust for machines among their users.","url_abs":"http://arxiv.org/abs/1612.00837v3","url_pdf":"http://arxiv.org/pdf/1612.00837v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"making-the-v-in-vqa-matter-elevating-the-role","repo_url":"https://github.com/SatyamGaba/visual_question_answering","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"making-the-v-in-vqa-matter-elevating-the-role","repo_url":"https://github.com/SatyamGaba/vqa","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"making-the-v-in-vqa-matter-elevating-the-role","repo_url":"https://github.com/abhshkdz/neural-vqa-attention","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"torch","reach":{"status":"ok"}},{"paper_slug":"making-the-v-in-vqa-matter-elevating-the-role","repo_url":"https://github.com/mokhalid-dev/Attention-based-VQA-model","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"making-the-v-in-vqa-matter-elevating-the-role","repo_url":"https://github.com/necla-ml/SNLI-VE","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"paper_slug":"making-the-v-in-vqa-matter-elevating-the-role","repo_url":"https://github.com/ntusteeian/VQA_CNN-LSTM","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"making-the-v-in-vqa-matter-elevating-the-role","repo_url":"https://github.com/yanxinyan1/yxy","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[{"slug":"visual-question-answering-v2-0","name":"Visual Question Answering v2.0","full_name":"VQA v2.0"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-question-answering-on-coco-visual","task":"Visual Question Answering (VQA)","dataset":"COCO Visual Question Answering (VQA) real images 2.0 open ended","model":"MCB","rank_in_archive_order":3,"of":4,"metrics":{"Percentage correct":"62.27"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-coco-visual","task":"Visual Question Answering (VQA)","dataset":"COCO Visual Question Answering (VQA) real images 2.0 open ended","model":"d-LSTM+nI","rank_in_archive_order":4,"of":4,"metrics":{"Percentage correct":"54.22"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-vqa-v2-test-std","task":"Visual Question Answering (VQA)","dataset":"VQA v2 test-std","model":"MCB [11, 12]","rank_in_archive_order":36,"of":38,"metrics":{"overall":"62.27"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-vqa-v2-test-std","task":"Visual Question Answering (VQA)","dataset":"VQA v2 test-std","model":"Language-only","rank_in_archive_order":37,"of":38,"metrics":{"overall":"44.26"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-vqa-v2-test-std","task":"Visual Question Answering (VQA)","dataset":"VQA v2 test-std","model":"Prior","rank_in_archive_order":38,"of":38,"metrics":{"overall":"25.98"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1612.00837","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}