Papers › Few-Shot Multimodal Explanation for Visual Question Answering

Few-Shot Multimodal Explanation for Visual Question Answering

28 Oct 2024ACM MM 2024 10archive 2025-07-28

Dizhan Xue, Shengsheng Qian, Changsheng Xu

A key object in eXplainable Artificial Intelligence (XAI) is to create intelligent systems capable of reasoning and explaining real-world data to facilitate reliable decision-making. Recent studies have acknowledged the importance of providing user-friendly and verifiable explanations to facilitate trustworthy Visual Question Answering (VQA) systems. This paper aims to promote explainable VQA from both data and method perspectives. First, we propose a new Standard Multimodal Explanation (SME) dataset and a new Few-Shot Multimodal Explanation for VQA (FS-MEVQA) task, which aims to generate the multimodal explanation of the underlying reasoning process for solving visual questions with few training samples. Our SME dataset includes 1,028,230 samples composed of questions, images, answers, and multimodal explanations, which can facilitate research in both traditional MEVQA and FS-MEVQA. To the best of our knowledge, this is the first large-scale dataset with joint language-vision explanations based on standard English and additional visual grounding tokens. Second, we propose a training-free Multimodal Explaining Agent (MEAgent) method based on an LLM agent with multimodal open-world tools to infer answers and generate multimodal explanations for visual questions. Our MEAgent can learn multimodal explanation from merely N(=16) training samples and leverage open-world abilities to perform FS-MEVQA on test samples. Comprehensive experimental results evaluated by language quality metrics, visual detection metric, and visual attribution metrics on our SME dataset indicate the superiority of our method for FS-MEVQA. Our code and data are available at https://github.com/LivXue/FS-MEVQA.

PaperPDFCode

Code

LivXue/FS-MEVQA mentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Explainable Artificial Intelligence (XAI)Explainable artificial intelligenceFS-MEVQAQuestion AnsweringVisual GroundingVisual Question AnsweringVisual Question Answering (VQA)

Datasets

Introduced by this paper, per the archive.

SME

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
FS-MEVQA SME MEAgent #Learning Samples (N) 16 #1 of 7 Archive leaderboard report
FS-MEVQA SME MEAgent ACC 51.45 #1 of 7 Archive leaderboard report
FS-MEVQA SME MEAgent BLEU-4 67.91 #1 of 7 Archive leaderboard report
FS-MEVQA SME MEAgent CIDEr 510.44 #1 of 7 Archive leaderboard report
FS-MEVQA SME MEAgent Detection 29.09 #1 of 7 Archive leaderboard report
FS-MEVQA SME MEAgent METEOR 50.55 #1 of 7 Archive leaderboard report
FS-MEVQA SME MEAgent ROUGE-L 79.41 #1 of 7 Archive leaderboard report
FS-MEVQA SME MEAgent SPICE 64.09 #1 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections