{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fvqa-fact-based-visual-question-answering","title":"FVQA: Fact-based Visual Question Answering","arxiv_id":"1606.05433","date":"2016-06-17","proceeding":null,"authors":["Peng Wang","Qi Wu","Chunhua Shen","Anton Van Den Hengel","Anthony Dick"],"abstract":"Visual Question Answering (VQA) has attracted a lot of attention in both\nComputer Vision and Natural Language Processing communities, not least because\nit offers insight into the relationships between two important sources of\ninformation. Current datasets, and the models built upon them, have focused on\nquestions which are answerable by direct analysis of the question and image\nalone. The set of such questions that require no external information to answer\nis interesting, but very limited. It excludes questions which require common\nsense, or basic factual knowledge to answer, for example. Here we introduce\nFVQA, a VQA dataset which requires, and supports, much deeper reasoning. FVQA\nonly contains questions which require external information to answer.\n  We thus extend a conventional visual question answering dataset, which\ncontains image-question-answerg triplets, through additional\nimage-question-answer-supporting fact tuples. The supporting fact is\nrepresented as a structural triplet, such as <Cat,CapableOf,ClimbingTrees>.\n  We evaluate several baseline models on the FVQA dataset, and describe a novel\nmodel which is capable of reasoning about an image on the basis of supporting\nfacts.","url_abs":"http://arxiv.org/abs/1606.05433v4","url_pdf":"http://arxiv.org/pdf/1606.05433v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"common-sense-reasoning","task_name":"Common Sense Reasoning"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":null,"task_name":"Triplet"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-question-answering-on-f-vqa","task":"Visual Question Answering (VQA)","dataset":"F-VQA","model":"F-VQA (top-3-QQmaping)","rank_in_archive_order":2,"of":3,"metrics":{"Top-1 Accuracy":"56.91","Top-3 Accuracy":"64.65"},"uses_additional_data":false},{"leaderboard":"/sota/visual-question-answering-on-f-vqa","task":"Visual Question Answering (VQA)","dataset":"F-VQA","model":"F-VQA (top-1-QQmaping)","rank_in_archive_order":3,"of":3,"metrics":{"Top-1 Accuracy":"52.56","Top-3 Accuracy":"59.72"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1606.05433","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}