{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/playpen-an-environment-for-exploring-learning","title":"Playpen: An Environment for Exploring Learning Through Conversational Interaction","arxiv_id":"2504.08590","date":"2025-04-11","proceeding":null,"authors":["Nicola Horst","Davide Mazzaccara","Antonia Schmidt","Michael Sullivan","Filippo Momentè","Luca Franceschetti","Philipp Sadler","Sherzod Hakimov","Alberto Testoni","Raffaella Bernardi","Raquel Fernández","Alexander Koller","Oliver Lemon","David Schlangen","Mario Giulianelli","Alessandro Suglia"],"abstract":"Are we running out of learning signal? Predicting the next word in an existing text has turned out to be a powerful signal, at least at scale. But there are signs that we are running out of this resource. In recent months, interaction between learner and feedback-giver has come into focus, both for \"alignment\" (with a reward model judging the quality of instruction following attempts) and for improving \"reasoning\" (process- and outcome-based verifiers judging reasoning steps). In this paper, we explore to what extent synthetic interaction in what we call Dialogue Games -- goal-directed and rule-governed activities driven predominantly by verbal actions -- can provide a learning signal, and how this signal can be used. We introduce an environment for producing such interaction data (with the help of a Large Language Model as counterpart to the learner model), both offline and online. We investigate the effects of supervised fine-tuning on this data, as well as reinforcement learning setups such as DPO, and GRPO; showing that all of these approaches achieve some improvements in in-domain games, but only GRPO demonstrates the ability to generalise to out-of-domain games as well as retain competitive performance in reference-based tasks. We release the framework and the baseline training setups in the hope that this can foster research in this promising new direction.","url_abs":"https://arxiv.org/abs/2504.08590v1","url_pdf":"https://arxiv.org/pdf/2504.08590v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"playpen-an-environment-for-exploring-learning","repo_url":"https://github.com/lm-playpen/playpen","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"instruction-following","task_name":"Instruction Following"},{"task_slug":"large-language-model","task_name":"Large Language Model"}],"methods":[{"method_slug":"dpo","method_name":"DPO"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}