{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/iqa-visual-question-answering-in-interactive","title":"IQA: Visual Question Answering in Interactive Environments","arxiv_id":"1712.03316","date":"2017-12-09","proceeding":"CVPR 2018 6","authors":["Daniel Gordon","Aniruddha Kembhavi","Mohammad Rastegari","Joseph Redmon","Dieter Fox","Ali Farhadi"],"abstract":"We introduce Interactive Question Answering (IQA), the task of answering\nquestions that require an autonomous agent to interact with a dynamic visual\nenvironment. IQA presents the agent with a scene and a question, like: \"Are\nthere any apples in the fridge?\" The agent must navigate around the scene,\nacquire visual understanding of scene elements, interact with objects (e.g.\nopen refrigerators) and plan for a series of actions conditioned on the\nquestion. Popular reinforcement learning approaches with a single controller\nperform poorly on IQA owing to the large and diverse state space. We propose\nthe Hierarchical Interactive Memory Network (HIMN), consisting of a factorized\nset of controllers, allowing the system to operate at multiple levels of\ntemporal abstraction. To evaluate HIMN, we introduce IQUAD V1, a new dataset\nbuilt upon AI2-THOR, a simulated photo-realistic environment of configurable\nindoor scenes with interactive objects (code and dataset available at\nhttps://github.com/danielgordon10/thor-iqa-cvpr-2018). IQUAD V1 has 75,000\nquestions, each paired with a unique scene configuration. Our experiments show\nthat our proposed model outperforms popular single controller based methods on\nIQUAD V1. For sample questions and results, please view our video:\nhttps://youtu.be/pXd3C-1jr98","url_abs":"http://arxiv.org/abs/1712.03316v3","url_pdf":"http://arxiv.org/pdf/1712.03316v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"iqa-visual-question-answering-in-interactive","repo_url":"https://github.com/danielgordon10/thor-iqa-cvpr-2018","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"navigate","task_name":"Navigate"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"},{"task_slug":"visual-question-answering","task_name":"Visual Question Answering (VQA)"}],"methods":[{"method_slug":"memory-network","method_name":"Memory Network"}],"datasets_introduced":[{"slug":"iquad","name":"IQUAD","full_name":"Interactive Question Answering Dataset"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1712.03316","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}