{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-promise-of-premise-harnessing-question","title":"The Promise of Premise: Harnessing Question Premises in Visual Question Answering","arxiv_id":"1705.00601","date":"2017-05-01","proceeding":"EMNLP 2017 9","authors":["Aroma Mahendru","Viraj Prabhu","Akrit Mohapatra","Dhruv Batra","Stefan Lee"],"abstract":"In this paper, we make a simple observation that questions about images often\ncontain premises - objects and relationships implied by the question - and that\nreasoning about premises can help Visual Question Answering (VQA) models\nrespond more intelligently to irrelevant or previously unseen questions. When\npresented with a question that is irrelevant to an image, state-of-the-art VQA\nmodels will still answer purely based on learned language biases, resulting in\nnon-sensical or even misleading answers. We note that a visual question is\nirrelevant to an image if at least one of its premises is false (i.e. not\ndepicted in the image). We leverage this observation to construct a dataset for\nQuestion Relevance Prediction and Explanation (QRPE) by searching for false\npremises. We train novel question relevance detection models and show that\nmodels that reason about premises consistently outperform models that do not.\nWe also find that forcing standard VQA models to reason about premises during\ntraining can lead to improvements on tasks requiring compositional reasoning.","url_abs":"http://arxiv.org/abs/1705.00601v2","url_pdf":"http://arxiv.org/pdf/1705.00601v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-promise-of-premise-harnessing-question","repo_url":"https://github.com/virajprabhu/premise-emnlp17","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"relevance-detection","task_name":"Relevance Detection"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1705.00601","atlas_url":"https://app.syntology.ai/?focus=1705.00601","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}