{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dialogue-learning-with-human-teaching-and","title":"Dialogue Learning with Human Teaching and Feedback in End-to-End Trainable Task-Oriented Dialogue Systems","arxiv_id":"1804.06512","date":"2018-04-18","proceeding":"NAACL 2018 6","authors":["Bing Liu","Gokhan Tur","Dilek Hakkani-Tur","Pararth Shah","Larry Heck"],"abstract":"In this work, we present a hybrid learning method for training task-oriented\ndialogue systems through online user interactions. Popular methods for learning\ntask-oriented dialogues include applying reinforcement learning with user\nfeedback on supervised pre-training models. Efficiency of such learning method\nmay suffer from the mismatch of dialogue state distribution between offline\ntraining and online interactive learning stages. To address this challenge, we\npropose a hybrid imitation and reinforcement learning method, with which a\ndialogue agent can effectively learn from its interaction with users by\nlearning from human teaching and feedback. We design a neural network based\ntask-oriented dialogue agent that can be optimized end-to-end with the proposed\nlearning method. Experimental results show that our end-to-end dialogue agent\ncan learn effectively from the mistake it makes via imitation learning from\nuser teaching. Applying reinforcement learning with user feedback after the\nimitation learning stage further improves the agent's capability in\nsuccessfully completing a task.","url_abs":"http://arxiv.org/abs/1804.06512v1","url_pdf":"http://arxiv.org/pdf/1804.06512v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dialogue-learning-with-human-teaching-and","repo_url":"https://github.com/google-research-datasets/simulated-dialogue","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"dialogue-state-tracking","task_name":"Dialogue State Tracking"},{"task_slug":"imitation-learning","task_name":"Imitation Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"task-oriented-dialogue-systems","task_name":"Task-Oriented Dialogue Systems"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/dialogue-state-tracking-on-second-dialogue","task":"Dialogue State Tracking","dataset":"Second dialogue state tracking challenge","model":"Liu et al.","rank_in_archive_order":6,"of":7,"metrics":{"Area":"90","Food":"84","Joint":"72","Price":"92","Request":"-"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.06512","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}