{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-many-random-seeds-statistical-power","title":"How Many Random Seeds? Statistical Power Analysis in Deep Reinforcement Learning Experiments","arxiv_id":"1806.08295","date":"2018-06-21","proceeding":null,"authors":["Cédric Colas","Olivier Sigaud","Pierre-Yves Oudeyer"],"abstract":"Consistently checking the statistical significance of experimental results is\none of the mandatory methodological steps to address the so-called\n\"reproducibility crisis\" in deep reinforcement learning. In this tutorial\npaper, we explain how the number of random seeds relates to the probabilities\nof statistical errors. For both the t-test and the bootstrap confidence\ninterval test, we recall theoretical guidelines to determine the number of\nrandom seeds one should use to provide a statistically significant comparison\nof the performance of two algorithms. Finally, we discuss the influence of\ndeviations from the assumptions usually made by statistical tests. We show that\nthey can lead to inaccurate evaluations of statistical errors and provide\nguidelines to counter these negative effects. We make our code available to\nperform the tests.","url_abs":"http://arxiv.org/abs/1806.08295v2","url_pdf":"http://arxiv.org/pdf/1806.08295v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-many-random-seeds-statistical-power","repo_url":"https://github.com/flowersteam/rl-difference-testing","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1806.08295","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}