{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/strategic-dialogue-management-via-deep","title":"Strategic Dialogue Management via Deep Reinforcement Learning","arxiv_id":"1511.08099","date":"2015-11-25","proceeding":null,"authors":["Heriberto Cuayáhuitl","Simon Keizer","Oliver Lemon"],"abstract":"Artificially intelligent agents equipped with strategic skills that can\nnegotiate during their interactions with other natural or artificial agents are\nstill underdeveloped. This paper describes a successful application of Deep\nReinforcement Learning (DRL) for training intelligent agents with strategic\nconversational skills, in a situated dialogue setting. Previous studies have\nmodelled the behaviour of strategic agents using supervised learning and\ntraditional reinforcement learning techniques, the latter using tabular\nrepresentations or learning with linear function approximation. In this study,\nwe apply DRL with a high-dimensional state space to the strategic board game of\nSettlers of Catan---where players can offer resources in exchange for others\nand they can also reply to offers made by other players. Our experimental\nresults report that the DRL-based learnt policies significantly outperformed\nseveral baselines including random, rule-based, and supervised-based\nbehaviours. The DRL-based policy has a 53% win rate versus 3 automated players\n(`bots'), whereas a supervised player trained on a dialogue corpus in this\nsetting achieved only 27%, versus the same 3 bots. This result supports the\nclaim that DRL is a promising framework for training dialogue systems, and\nstrategic agents with negotiation abilities.","url_abs":"http://arxiv.org/abs/1511.08099v1","url_pdf":"http://arxiv.org/pdf/1511.08099v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"strategic-dialogue-management-via-deep","repo_url":"https://github.com/cuayahuitl/SimpleDS","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"dialogue-management","task_name":"Dialogue Management"},{"task_slug":"management","task_name":"Management"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1511.08099","atlas_url":"https://app.syntology.ai/?focus=1511.08099","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}