Papers › Variational Causal Inference Network for Explanatory Visual Question Answering

Variational Causal Inference Network for Explanatory Visual Question Answering

1 Jan 2023ICCV 2023 1archive 2025-07-28

Dizhan Xue, Shengsheng Qian, Changsheng Xu

Explanatory Visual Question Answering (EVQA) is a recently proposed multimodal reasoning task that requires answering visual questions and generating multimodal explanations for the reasoning processes. Unlike traditional Visual Question Answering (VQA) which focuses solely on answering, EVQA aims to provide user-friendly explanations to enhance the explainability and credibility of reasoning models. However, existing EVQA methods typically predict the answer and explanation separately, which ignores the causal correlation between them. Moreover, they neglect the complex relationships among question words, visual regions, and explanation tokens. To address these issues, we propose a Variational Causal Inference Network (VCIN) that establishes the causal correlation between predicted answers and explanations, and captures cross-modal relationships to generate rational explanations. First, we utilize a vision-and-language pretrained model to extract visual features and question features. Secondly, we propose a multimodal explanation gating transformer that constructs cross-modal relationships and generates rational explanations. Finally, we propose a variational causal inference to establish the target causal structure and predict the answers. Comprehensive experiments demonstrate the superiority of VCIN over state-of-the-art EVQA methods.

PaperPDFCode

Code

LivXue/VCIN officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Explanation GenerationExplanatory Visual Question AnsweringFS-MEVQAMultimodal ReasoningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Explanatory Visual Question Answering GQA-REX VCIN BLEU-4 58.65 #1 of 5 Archive leaderboard report
Explanatory Visual Question Answering GQA-REX VCIN CIDEr 519.23 #1 of 5 Archive leaderboard report
Explanatory Visual Question Answering GQA-REX VCIN GQA-test 60.61 #1 of 5 Archive leaderboard report
Explanatory Visual Question Answering GQA-REX VCIN GQA-val 81.80 #1 of 5 Archive leaderboard report
Explanatory Visual Question Answering GQA-REX VCIN Grounding 77.33 #1 of 5 Archive leaderboard report
Explanatory Visual Question Answering GQA-REX VCIN METEOR 41.57 #1 of 5 Archive leaderboard report
Explanatory Visual Question Answering GQA-REX VCIN ROUGE-L 81.45 #1 of 5 Archive leaderboard report
Explanatory Visual Question Answering GQA-REX VCIN SPICE 54.63 #1 of 5 Archive leaderboard report
FS-MEVQA SME VCIN #Learning Samples (N) 16 #6 of 7 Archive leaderboard report
FS-MEVQA SME VCIN ACC 17.77 #6 of 7 Archive leaderboard report
FS-MEVQA SME VCIN BLEU-4 9.17 #6 of 7 Archive leaderboard report
FS-MEVQA SME VCIN CIDEr 4.28 #6 of 7 Archive leaderboard report
FS-MEVQA SME VCIN Detection 0.28 #6 of 7 Archive leaderboard report
FS-MEVQA SME VCIN METEOR 19.82 #6 of 7 Archive leaderboard report
FS-MEVQA SME VCIN ROUGE-L 33.34 #6 of 7 Archive leaderboard report
FS-MEVQA SME VCIN SPICE 13.39 #6 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Causal inference

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections