{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mapping-instructions-to-actions-in-3d","title":"Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction","arxiv_id":"1809.00786","date":"2018-09-04","proceeding":"EMNLP 2018 10","authors":["Dipendra Misra","Andrew Bennett","Valts Blukis","Eyvind Niklasson","Max Shatkhin","Yoav Artzi"],"abstract":"We propose to decompose instruction execution to goal prediction and action\ngeneration. We design a model that maps raw visual observations to goals using\nLINGUNET, a language-conditioned image generation network, and then generates\nthe actions required to complete them. Our model is trained from demonstration\nonly without external resources. To evaluate our approach, we introduce two\nbenchmarks for instruction following: LANI, a navigation task; and CHAI, where\nan agent executes household instructions. Our evaluation demonstrates the\nadvantages of our model decomposition, and illustrates the challenges posed by\nour new benchmarks.","url_abs":"http://arxiv.org/abs/1809.00786v2","url_pdf":"http://arxiv.org/pdf/1809.00786v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mapping-instructions-to-actions-in-3d","repo_url":"https://github.com/clic-lab/ciff","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"mapping-instructions-to-actions-in-3d","repo_url":"https://github.com/clic-lab/chalet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"mapping-instructions-to-actions-in-3d","repo_url":"https://github.com/clic-lab/drif","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"mapping-instructions-to-actions-in-3d","repo_url":"https://github.com/lil-lab/chalet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"mapping-instructions-to-actions-in-3d","repo_url":"https://github.com/lil-lab/ciff","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"action-generation","task_name":"Action Generation"},{"task_slug":"conditional-image-generation","task_name":"Conditional Image Generation"},{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"instruction-following","task_name":"Instruction Following"}],"methods":[],"datasets_introduced":[{"slug":"lani","name":"Lani","full_name":null}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1809.00786","atlas_url":"https://app.syntology.ai/?focus=1809.00786","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}