{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/whats-in-a-question-using-visual-questions-as","title":"What's in a Question: Using Visual Questions as a Form of Supervision","arxiv_id":"1704.03895","date":"2017-04-12","proceeding":"CVPR 2017 7","authors":["Siddha Ganju","Olga Russakovsky","Abhinav Gupta"],"abstract":"Collecting fully annotated image datasets is challenging and expensive. Many\ntypes of weak supervision have been explored: weak manual annotations, web\nsearch results, temporal continuity, ambient sound and others. We focus on one\nparticular unexplored mode: visual questions that are asked about images. The\nkey observation that inspires our work is that the question itself provides\nuseful information about the image (even without the answer being available).\nFor instance, the question \"what is the breed of the dog?\" informs the AI that\nthe animal in the scene is a dog and that there is only one dog present. We\nmake three contributions: (1) providing an extensive qualitative and\nquantitative analysis of the information contained in human visual questions,\n(2) proposing two simple but surprisingly effective modifications to the\nstandard visual question answering models that allow them to make use of weak\nsupervision in the form of unanswered questions associated with images and (3)\ndemonstrating that a simple data augmentation strategy inspired by our insights\nresults in a 7.1% improvement on the standard VQA benchmark.","url_abs":"http://arxiv.org/abs/1704.03895v1","url_pdf":"http://arxiv.org/pdf/1704.03895v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"whats-in-a-question-using-visual-questions-as","repo_url":"https://github.com/sidgan/whats_in_a_question","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"form","task_name":"Form"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.03895","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}