{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/revisiting-visual-question-answering","title":"Revisiting Visual Question Answering Baselines","arxiv_id":"1606.08390","date":"2016-06-27","proceeding":null,"authors":["Allan Jabri","Armand Joulin","Laurens van der Maaten"],"abstract":"Visual question answering (VQA) is an interesting learning setting for\nevaluating the abilities and shortcomings of current systems for image\nunderstanding. Many of the recently proposed VQA systems include attention or\nmemory mechanisms designed to support \"reasoning\". For multiple-choice VQA,\nnearly all of these systems train a multi-class classifier on image and\nquestion features to predict an answer. This paper questions the value of these\ncommon practices and develops a simple alternative model based on binary\nclassification. Instead of treating answers as competing choices, our model\nreceives the answer as input and predicts whether or not an\nimage-question-answer triplet is correct. We evaluate our model on the Visual7W\nTelling and the VQA Real Multiple Choice tasks, and find that even simple\nversions of our model perform competitively. Our best model achieves\nstate-of-the-art performance on the Visual7W Telling task and compares\nsurprisingly well with the most complex systems proposed for the VQA Real\nMultiple Choice task. We explore variants of the model and study its\ntransferability between both datasets. We also present an error analysis of our\nmodel that suggests a key problem of current VQA systems lies in the lack of\nvisual grounding of concepts that occur in the questions and answers. Overall,\nour results suggest that the performance of current VQA systems is not\nsignificantly better than that of systems designed to exploit dataset biases.","url_abs":"http://arxiv.org/abs/1606.08390v2","url_pdf":"http://arxiv.org/pdf/1606.08390v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"revisiting-visual-question-answering","repo_url":"https://github.com/Cold-Winter/vqs","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"caffe2","reach":{"status":"unanswered"}},{"paper_slug":"revisiting-visual-question-answering","repo_url":"https://github.com/arunmallya/simple-vqa","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"torch","reach":{"status":"ok"}},{"paper_slug":"revisiting-visual-question-answering","repo_url":"https://github.com/zixuwang1996/VQA-reading-list","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"binary-classification","task_name":"Binary Classification"},{"task_slug":"multiple-choice","task_name":"Multiple-choice"},{"task_slug":null,"task_name":"Triplet"},{"task_slug":"visual-grounding","task_name":"Visual Grounding"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}