{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dialogue-generation-from-imitation-learning","title":"Dialogue Generation: From Imitation Learning to Inverse Reinforcement Learning","arxiv_id":"1812.03509","date":"2018-12-09","proceeding":null,"authors":["Ziming Li","Julia Kiseleva","Maarten de Rijke"],"abstract":"The performance of adversarial dialogue generation models relies on the\nquality of the reward signal produced by the discriminator. The reward signal\nfrom a poor discriminator can be very sparse and unstable, which may lead the\ngenerator to fall into a local optimum or to produce nonsense replies. To\nalleviate the first problem, we first extend a recently proposed adversarial\ndialogue generation method to an adversarial imitation learning solution. Then,\nin the framework of adversarial inverse reinforcement learning, we propose a\nnew reward model for dialogue generation that can provide a more accurate and\nprecise reward signal for generator training. We evaluate the performance of\nthe resulting model with automatic metrics and human evaluations in two\nannotation settings. Our experimental results demonstrate that our model can\ngenerate more high-quality responses and achieve higher overall performance\nthan the state-of-the-art.","url_abs":"http://arxiv.org/abs/1812.03509v1","url_pdf":"http://arxiv.org/pdf/1812.03509v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dialogue-generation-from-imitation-learning","repo_url":"https://bitbucket.org/ZimingLi/dg-irl-aaai2019","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"dialogue-generation","task_name":"Dialogue Generation"},{"task_slug":"imitation-learning","task_name":"Imitation Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.03509","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}