{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lets-play-again-variability-of-deep","title":"Let's Play Again: Variability of Deep Reinforcement Learning Agents in Atari Environments","arxiv_id":"1904.06312","date":"2019-04-12","proceeding":null,"authors":["Kaleigh Clary","Emma Tosch","John Foley","David Jensen"],"abstract":"Reproducibility in reinforcement learning is challenging: uncontrolled\nstochasticity from many sources, such as the learning algorithm, the learned\npolicy, and the environment itself have led researchers to report the\nperformance of learned agents using aggregate metrics of performance over\nmultiple random seeds for a single environment. Unfortunately, there are still\npernicious sources of variability in reinforcement learning agents that make\nreporting common summary statistics an unsound metric for performance. Our\nexperiments demonstrate the variability of common agents used in the popular\nOpenAI Baselines repository. We make the case for reporting post-training agent\nperformance as a distribution, rather than a point estimate.","url_abs":"http://arxiv.org/abs/1904.06312v1","url_pdf":"http://arxiv.org/pdf/1904.06312v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lets-play-again-variability-of-deep","repo_url":"https://github.com/kclary/variability-RL","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"atari-games","task_name":"Atari Games"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1904.06312","atlas_url":"https://app.syntology.ai/?focus=1904.06312","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}