{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ask-your-neurons-a-deep-learning-approach-to","title":"Ask Your Neurons: A Deep Learning Approach to Visual Question Answering","arxiv_id":"1605.02697","date":"2016-05-09","proceeding":null,"authors":["Mateusz Malinowski","Marcus Rohrbach","Mario Fritz"],"abstract":"We address a question answering task on real-world images that is set up as a\nVisual Turing Test. By combining latest advances in image representation and\nnatural language processing, we propose Ask Your Neurons, a scalable, jointly\ntrained, end-to-end formulation to this problem.\n  In contrast to previous efforts, we are facing a multi-modal problem where\nthe language output (answer) is conditioned on visual and natural language\ninputs (image and question). We provide additional insights into the problem by\nanalyzing how much information is contained only in the language part for which\nwe provide a new human baseline. To study human consensus, which is related to\nthe ambiguities inherent in this challenging task, we propose two novel metrics\nand collect additional answers which extend the original DAQUAR dataset to\nDAQUAR-Consensus.\n  Moreover, we also extend our analysis to VQA, a large-scale question\nanswering about images dataset, where we investigate some particular design\nchoices and show the importance of stronger visual models. At the same time, we\nachieve strong performance of our model that still uses a global image\nrepresentation. Finally, based on such analysis, we refine our Ask Your Neurons\non DAQUAR, which also leads to a better performance on this challenging task.","url_abs":"http://arxiv.org/abs/1605.02697v2","url_pdf":"http://arxiv.org/pdf/1605.02697v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ask-your-neurons-a-deep-learning-approach-to","repo_url":"https://github.com/mateuszmalinowski/visual_turing_test-tutorial","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}