{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mapping-instructions-and-visual-observations","title":"Mapping Instructions and Visual Observations to Actions with Reinforcement Learning","arxiv_id":"1704.08795","date":"2017-04-28","proceeding":"EMNLP 2017 9","authors":["Dipendra Misra","John Langford","Yoav Artzi"],"abstract":"We propose to directly map raw visual observations and text input to actions\nfor instruction execution. While existing approaches assume access to\nstructured environment representations or use a pipeline of separately trained\nmodels, we learn a single model to jointly reason about linguistic and visual\ninput. We use reinforcement learning in a contextual bandit setting to train a\nneural network agent. To guide the agent's exploration, we use reward shaping\nwith different forms of supervision. Our approach does not require intermediate\nrepresentations, planning procedures, or training different models. We evaluate\nin a simulated environment, and show significant improvements over supervised\nlearning and common reinforcement learning variants.","url_abs":"http://arxiv.org/abs/1704.08795v2","url_pdf":"http://arxiv.org/pdf/1704.08795v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mapping-instructions-and-visual-observations","repo_url":"https://github.com/clic-lab/blocks","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.08795","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}