{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/analyzing-the-behavior-of-visual-question","title":"Analyzing the Behavior of Visual Question Answering Models","arxiv_id":"1606.07356","date":"2016-06-23","proceeding":"EMNLP 2016 11","authors":["Aishwarya Agrawal","Dhruv Batra","Devi Parikh"],"abstract":"Recently, a number of deep-learning based models have been proposed for the\ntask of Visual Question Answering (VQA). The performance of most models is\nclustered around 60-70%. In this paper we propose systematic methods to analyze\nthe behavior of these models as a first step towards recognizing their\nstrengths and weaknesses, and identifying the most fruitful directions for\nprogress. We analyze two models, one each from two major classes of VQA models\n-- with-attention and without-attention and show the similarities and\ndifferences in the behavior of these models. We also analyze the winning entry\nof the VQA Challenge 2016.\n  Our behavior analysis reveals that despite recent progress, today's VQA\nmodels are \"myopic\" (tend to fail on sufficiently novel instances), often \"jump\nto conclusions\" (converge on a predicted answer after 'listening' to just half\nthe question), and are \"stubborn\" (do not change their answers across images).","url_abs":"http://arxiv.org/abs/1606.07356v2","url_pdf":"http://arxiv.org/pdf/1606.07356v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"analyzing-the-behavior-of-visual-question","repo_url":"https://github.com/akirafukui/vqa-mcb","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"caffe2","reach":{"status":"ok","spdx":"BSD-2-Clause"}}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1606.07356","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1606.07356"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/akirafukui/vqa-mcb","reach":{"status":"ok","spdx":"BSD-2-Clause"}}],"summary":{"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"fe0cc787192a0db6","entry":"remove_words","repo":"akirafukui/vqa-mcb","repo_kind":"official","path":"preprocess/vg_preprocessing.py","file_url":"https://github.com/akirafukui/vqa-mcb/blob/HEAD/preprocess/vg_preprocessing.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-2-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"fe0cc787192a0db6"}},{"code_sha256_prefix":"37aeb488bf9c3c2d","entry":"text2int","repo":"akirafukui/vqa-mcb","repo_kind":"official","path":"preprocess/vg_preprocessing.py","file_url":"https://github.com/akirafukui/vqa-mcb/blob/HEAD/preprocess/vg_preprocessing.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-2-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"37aeb488bf9c3c2d"}},{"code_sha256_prefix":"e2c645c25e14ea4b","entry":"tokenize","repo":"akirafukui/vqa-mcb","repo_kind":"official","path":"preprocess/vg_preprocessing.py","file_url":"https://github.com/akirafukui/vqa-mcb/blob/HEAD/preprocess/vg_preprocessing.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-2-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"e2c645c25e14ea4b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}