{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-end-to-end-reinforcement-learning-of","title":"Towards End-to-End Reinforcement Learning of Dialogue Agents for Information Access","arxiv_id":"1609.00777","date":"2016-09-03","proceeding":"ACL 2017 7","authors":["Bhuwan Dhingra","Lihong Li","Xiujun Li","Jianfeng Gao","Yun-Nung Chen","Faisal Ahmed","Li Deng"],"abstract":"This paper proposes KB-InfoBot -- a multi-turn dialogue agent which helps\nusers search Knowledge Bases (KBs) without composing complicated queries. Such\ngoal-oriented dialogue agents typically need to interact with an external\ndatabase to access real-world knowledge. Previous systems achieved this by\nissuing a symbolic query to the KB to retrieve entries based on their\nattributes. However, such symbolic operations break the differentiability of\nthe system and prevent end-to-end training of neural dialogue agents. In this\npaper, we address this limitation by replacing symbolic queries with an induced\n\"soft\" posterior distribution over the KB that indicates which entities the\nuser is interested in. Integrating the soft retrieval process with a\nreinforcement learner leads to higher task success rate and reward in both\nsimulations and against real users. We also present a fully neural end-to-end\nagent, trained entirely from user feedback, and discuss its application towards\npersonalized dialogue agents. The source code is available at\nhttps://github.com/MiuLab/KB-InfoBot.","url_abs":"http://arxiv.org/abs/1609.00777v3","url_pdf":"http://arxiv.org/pdf/1609.00777v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-end-to-end-reinforcement-learning-of","repo_url":"https://github.com/MiuLab/KB-InfoBot","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"world-knowledge","task_name":"World Knowledge"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1609.00777","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}