{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/active-learning-for-visual-question-answering","title":"Active Learning for Visual Question Answering: An Empirical Study","arxiv_id":"1711.01732","date":"2017-11-06","proceeding":null,"authors":["Xiao Lin","Devi Parikh"],"abstract":"We present an empirical study of active learning for Visual Question\nAnswering, where a deep VQA model selects informative question-image pairs from\na pool and queries an oracle for answers to maximally improve its performance\nunder a limited query budget. Drawing analogies from human learning, we explore\ncramming (entropy), curiosity-driven (expected model change), and goal-driven\n(expected error reduction) active learning approaches, and propose a fast and\neffective goal-driven active learning scoring function to pick question-image\npairs for deep VQA models under the Bayesian Neural Network framework. We find\nthat deep VQA models need large amounts of training data before they can start\nasking informative questions. But once they do, all three approaches outperform\nthe random selection baseline and achieve significant query savings. For the\nscenario where the model is allowed to ask generic questions about images but\nis evaluated only on specific questions (e.g., questions whose answer is either\nyes or no), our proposed goal-driven scoring function performs the best.","url_abs":"http://arxiv.org/abs/1711.01732v1","url_pdf":"http://arxiv.org/pdf/1711.01732v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"active-learning-for-visual-question-answering","repo_url":"https://github.com/frkl/active-learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"torch","reach":null}],"tasks":[{"task_slug":"active-learning","task_name":"Active Learning"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1711.01732","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}