{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/answer-them-all-toward-universal-visual","title":"Answer Them All! Toward Universal Visual Question Answering Models","arxiv_id":"1903.00366","date":"2019-03-01","proceeding":"CVPR 2019 6","authors":["Robik Shrestha","Kushal Kafle","Christopher Kanan"],"abstract":"Visual Question Answering (VQA) research is split into two camps: the first\nfocuses on VQA datasets that require natural image understanding and the second\nfocuses on synthetic datasets that test reasoning. A good VQA algorithm should\nbe capable of both, but only a few VQA algorithms are tested in this manner. We\ncompare five state-of-the-art VQA algorithms across eight VQA datasets covering\nboth domains. To make the comparison fair, all of the models are standardized\nas much as possible, e.g., they use the same visual features, answer\nvocabularies, etc. We find that methods do not generalize across the two\ndomains. To address this problem, we propose a new VQA algorithm that rivals or\nexceeds the state-of-the-art for both domains.","url_abs":"http://arxiv.org/abs/1903.00366v2","url_pdf":"http://arxiv.org/pdf/1903.00366v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"answer-them-all-toward-universal-visual","repo_url":"https://github.com/erobic/ramen","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"answer-them-all-toward-universal-visual","repo_url":"https://github.com/bhanukamanesha/ramen","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"all","task_name":"All"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.00366","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}