Papers › FVQA: Fact-based Visual Question Answering
FVQA: Fact-based Visual Question Answering
Peng Wang, Qi Wu, Chunhua Shen, Anton Van Den Hengel, Anthony Dick
Visual Question Answering (VQA) has attracted a lot of attention in both Computer Vision and Natural Language Processing communities, not least because it offers insight into the relationships between two important sources of information. Current datasets, and the models built upon them, have focused on questions which are answerable by direct analysis of the question and image alone. The set of such questions that require no external information to answer is interesting, but very limited. It excludes questions which require common sense, or basic factual knowledge to answer, for example. Here we introduce FVQA, a VQA dataset which requires, and supports, much deeper reasoning. FVQA only contains questions which require external information to answer. We thus extend a conventional visual question answering dataset, which contains image-question-answerg triplets, through additional image-question-answer-supporting fact tuples. The supporting fact is represented as a structural triplet, such as <Cat,CapableOf,ClimbingTrees>. We evaluate several baseline models on the FVQA dataset, and describe a novel model which is capable of reasoning about an image on the basis of supporting facts.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Question Answering (VQA) | F-VQA | F-VQA (top-3-QQmaping) | Top-1 Accuracy | 56.91 | #2 of 3 | Archive leaderboard | report |
| Visual Question Answering (VQA) | F-VQA | F-VQA (top-3-QQmaping) | Top-3 Accuracy | 64.65 | #2 of 3 | Archive leaderboard | report |
| Visual Question Answering (VQA) | F-VQA | F-VQA (top-1-QQmaping) | Top-1 Accuracy | 52.56 | #3 of 3 | Archive leaderboard | report |
| Visual Question Answering (VQA) | F-VQA | F-VQA (top-1-QQmaping) | Top-3 Accuracy | 59.72 | #3 of 3 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections