{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-question-answering-a-survey-of-methods","title":"Visual Question Answering: A Survey of Methods and Datasets","arxiv_id":"1607.05910","date":"2016-07-20","proceeding":null,"authors":["Qi Wu","Damien Teney","Peng Wang","Chunhua Shen","Anthony Dick","Anton Van Den Hengel"],"abstract":"Visual Question Answering (VQA) is a challenging task that has received\nincreasing attention from both the computer vision and the natural language\nprocessing communities. Given an image and a question in natural language, it\nrequires reasoning over visual elements of the image and general knowledge to\ninfer the correct answer. In the first part of this survey, we examine the\nstate of the art by comparing modern approaches to the problem. We classify\nmethods by their mechanism to connect the visual and textual modalities. In\nparticular, we examine the common approach of combining convolutional and\nrecurrent neural networks to map images and questions to a common feature\nspace. We also discuss memory-augmented and modular architectures that\ninterface with structured knowledge bases. In the second part of this survey,\nwe review the datasets available for training and evaluating VQA systems. The\nvarious datatsets contain questions at different levels of complexity, which\nrequire different capabilities and types of reasoning. We examine in depth the\nquestion/answer pairs from the Visual Genome project, and evaluate the\nrelevance of the structured annotations of images with scene graphs for VQA.\nFinally, we discuss promising future directions for the field, in particular\nthe connection to structured knowledge bases and the use of natural language\nprocessing models.","url_abs":"http://arxiv.org/abs/1607.05910v1","url_pdf":"http://arxiv.org/pdf/1607.05910v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visual-question-answering-a-survey-of-methods","repo_url":"https://github.com/AI-metrics/AI-metrics","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"general-knowledge","task_name":"General Knowledge"},{"task_slug":"survey","task_name":"Survey"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1607.05910","atlas_url":"https://app.syntology.ai/?focus=1607.05910","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}