Papers › VQA-E: Explaining, Elaborating, and Enhancing Your Answers for Visual Questions
VQA-E: Explaining, Elaborating, and Enhancing Your Answers for Visual Questions
Qing Li, Qingyi Tao, Shafiq Joty, Jianfei Cai, Jiebo Luo
Most existing works in visual question answering (VQA) are dedicated to improving the accuracy of predicted answers, while disregarding the explanations. We argue that the explanation for an answer is of the same or even more importance compared with the answer itself, since it makes the question and answering process more understandable and traceable. To this end, we propose a new task of VQA-E (VQA with Explanation), where the computational models are required to generate an explanation with the predicted answer. We first construct a new dataset, and then frame the VQA-E problem in a multi-task learning architecture. Our VQA-E dataset is automatically derived from the VQA v2 dataset by intelligently exploiting the available captions. We have conducted a user study to validate the quality of explanations synthesized by our method. We quantitatively show that the additional supervision from explanations can not only produce insightful textual sentences to justify the answers, but also improve the performance of answer prediction. Our model outperforms the state-of-the-art methods by a clear margin on the VQA v2 dataset.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Datasets
Introduced by this paper, per the archive.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Explanatory Visual Question Answering | GQA-REX | VQAE | BLEU-4 | 42.56 | #4 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | VQAE | CIDEr | 358.20 | #4 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | VQAE | GQA-test | 57.24 | #4 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | VQAE | GQA-val | 65.19 | #4 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | VQAE | Grounding | 31.29 | #4 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | VQAE | METEOR | 34.51 | #4 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | VQAE | ROUGE-L | 73.59 | #4 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | VQAE | SPICE | 40.39 | #4 of 5 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections