Papers › FVQA: Fact-based Visual Question Answering

FVQA: Fact-based Visual Question Answering

17 Jun 2016arXiv:1606.05433archive 2025-07-28

Peng Wang, Qi Wu, Chunhua Shen, Anton Van Den Hengel, Anthony Dick

Visual Question Answering (VQA) has attracted a lot of attention in both Computer Vision and Natural Language Processing communities, not least because it offers insight into the relationships between two important sources of information. Current datasets, and the models built upon them, have focused on questions which are answerable by direct analysis of the question and image alone. The set of such questions that require no external information to answer is interesting, but very limited. It excludes questions which require common sense, or basic factual knowledge to answer, for example. Here we introduce FVQA, a VQA dataset which requires, and supports, much deeper reasoning. FVQA only contains questions which require external information to answer. We thus extend a conventional visual question answering dataset, which contains image-question-answerg triplets, through additional image-question-answer-supporting fact tuples. The supporting fact is represented as a structural triplet, such as <Cat,CapableOf,ClimbingTrees>. We evaluate several baseline models on the FVQA dataset, and describe a novel model which is capable of reasoning about an image on the basis of supporting facts.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Common Sense ReasoningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Question Answering (VQA) F-VQA F-VQA (top-3-QQmaping) Top-1 Accuracy 56.91 #2 of 3 Archive leaderboard report
Visual Question Answering (VQA) F-VQA F-VQA (top-3-QQmaping) Top-3 Accuracy 64.65 #2 of 3 Archive leaderboard report
Visual Question Answering (VQA) F-VQA F-VQA (top-1-QQmaping) Top-1 Accuracy 52.56 #3 of 3 Archive leaderboard report
Visual Question Answering (VQA) F-VQA F-VQA (top-1-QQmaping) Top-3 Accuracy 59.72 #3 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections