{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/reinforced-cross-modal-matching-and-self","title":"Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation","arxiv_id":"1811.10092","date":"2018-11-25","proceeding":"CVPR 2019 6","authors":["Xin Wang","Qiuyuan Huang","Asli Celikyilmaz","Jianfeng Gao","Dinghan Shen","Yuan-Fang Wang","William Yang Wang","Lei Zhang"],"abstract":"Vision-language navigation (VLN) is the task of navigating an embodied agent\nto carry out natural language instructions inside real 3D environments. In this\npaper, we study how to address three critical challenges for this task: the\ncross-modal grounding, the ill-posed feedback, and the generalization problems.\nFirst, we propose a novel Reinforced Cross-Modal Matching (RCM) approach that\nenforces cross-modal grounding both locally and globally via reinforcement\nlearning (RL). Particularly, a matching critic is used to provide an intrinsic\nreward to encourage global matching between instructions and trajectories, and\na reasoning navigator is employed to perform cross-modal grounding in the local\nvisual scene. Evaluation on a VLN benchmark dataset shows that our RCM model\nsignificantly outperforms previous methods by 10% on SPL and achieves the new\nstate-of-the-art performance. To improve the generalizability of the learned\npolicy, we further introduce a Self-Supervised Imitation Learning (SIL) method\nto explore unseen environments by imitating its own past, good decisions. We\ndemonstrate that SIL can approximate a better and more efficient policy, which\ntremendously minimizes the success rate performance gap between seen and unseen\nenvironments (from 30.7% to 11.7%).","url_abs":"http://arxiv.org/abs/1811.10092v2","url_pdf":"http://arxiv.org/pdf/1811.10092v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"imitation-learning","task_name":"Imitation Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"vision-language-navigation","task_name":"Vision-Language Navigation"},{"task_slug":"visual-navigation","task_name":"Visual Navigation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/vision-language-navigation-on-room2room","task":"Vision-Language Navigation","dataset":"Room2Room","model":"RCM + SIL","rank_in_archive_order":2,"of":3,"metrics":{"spl":"0.59"},"uses_additional_data":false},{"leaderboard":"/sota/visual-navigation-on-room-to-room-1","task":"Visual Navigation","dataset":"R2R","model":"RCM+SIL(no early exploration)","rank_in_archive_order":10,"of":11,"metrics":{"spl":"0.38"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1811.10092","atlas_url":"https://app.syntology.ai/?focus=1811.10092","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}