{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vqa-e-explaining-elaborating-and-enhancing","title":"VQA-E: Explaining, Elaborating, and Enhancing Your Answers for Visual Questions","arxiv_id":"1803.07464","date":"2018-03-20","proceeding":"ECCV 2018 9","authors":["Qing Li","Qingyi Tao","Shafiq Joty","Jianfei Cai","Jiebo Luo"],"abstract":"Most existing works in visual question answering (VQA) are dedicated to\nimproving the accuracy of predicted answers, while disregarding the\nexplanations. We argue that the explanation for an answer is of the same or\neven more importance compared with the answer itself, since it makes the\nquestion and answering process more understandable and traceable. To this end,\nwe propose a new task of VQA-E (VQA with Explanation), where the computational\nmodels are required to generate an explanation with the predicted answer. We\nfirst construct a new dataset, and then frame the VQA-E problem in a multi-task\nlearning architecture. Our VQA-E dataset is automatically derived from the VQA\nv2 dataset by intelligently exploiting the available captions. We have\nconducted a user study to validate the quality of explanations synthesized by\nour method. We quantitatively show that the additional supervision from\nexplanations can not only produce insightful textual sentences to justify the\nanswers, but also improve the performance of answer prediction. Our model\noutperforms the state-of-the-art methods by a clear margin on the VQA v2\ndataset.","url_abs":"http://arxiv.org/abs/1803.07464v2","url_pdf":"http://arxiv.org/pdf/1803.07464v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"explanatory-visual-question-answering","task_name":"Explanatory Visual Question Answering"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[{"slug":"vqa-e","name":"VQA-E","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/explanatory-visual-question-answering-on-gqa","task":"Explanatory Visual Question Answering","dataset":"GQA-REX","model":"VQAE","rank_in_archive_order":4,"of":5,"metrics":{"BLEU-4":"42.56","CIDEr":"358.20","GQA-test":"57.24","GQA-val":"65.19","Grounding":"31.29","METEOR":"34.51","ROUGE-L":"73.59","SPICE":"40.39"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1803.07464","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}