Papers › Faithful Multimodal Explanation for Visual Question Answering
Faithful Multimodal Explanation for Visual Question Answering
Jialin Wu, Raymond J. Mooney
AI systems' ability to explain their reasoning is critical to their utility and trustworthiness. Deep neural networks have enabled significant progress on many challenging problems such as visual question answering (VQA). However, most of them are opaque black boxes with limited explanatory capability. This paper presents a novel approach to developing a high-performing VQA system that can elucidate its answers with integrated textual and visual explanations that faithfully reflect important aspects of its underlying reasoning while capturing the style of comprehensible human explanations. Extensive experimental evaluation demonstrates the advantages of this approach compared to competing methods with both automatic evaluation metrics and human evaluation metrics.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Explanatory Visual Question Answering | GQA-REX | EXP | BLEU-4 | 42.45 | #5 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | EXP | CIDEr | 357.10 | #5 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | EXP | GQA-test | 56.92 | #5 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | EXP | GQA-val | 65.17 | #5 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | EXP | Grounding | 33.52 | #5 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | EXP | METEOR | 34.46 | #5 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | EXP | ROUGE-L | 73.51 | #5 of 5 | Archive leaderboard | report |
| Explanatory Visual Question Answering | GQA-REX | EXP | SPICE | 40.35 | #5 of 5 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections