{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/training-millions-of-personalized-dialogue","title":"Training Millions of Personalized Dialogue Agents","arxiv_id":"1809.01984","date":"2018-09-06","proceeding":"EMNLP 2018 10","authors":["Pierre-Emmanuel Mazaré","Samuel Humeau","Martin Raison","Antoine Bordes"],"abstract":"Current dialogue systems are not very engaging for users, especially when\ntrained end-to-end without relying on proactive reengaging scripted strategies.\nZhang et al. (2018) showed that the engagement level of end-to-end dialogue\nmodels increases when conditioning them on text personas providing some\npersonalized back-story to the model. However, the dataset used in Zhang et al.\n(2018) is synthetic and of limited size as it contains around 1k different\npersonas. In this paper we introduce a new dataset providing 5 million personas\nand 700 million persona-based dialogues. Our experiments show that, at this\nscale, training using personas still improves the performance of end-to-end\nsystems. In addition, we show that other tasks benefit from the wide coverage\nof our dataset by fine-tuning our model on the data from Zhang et al. (2018)\nand achieving state-of-the-art results.","url_abs":"http://arxiv.org/abs/1809.01984v1","url_pdf":"http://arxiv.org/pdf/1809.01984v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"training-millions-of-personalized-dialogue","repo_url":"https://github.com/Exe-dev/PersonaGeneration","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1809.01984","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}