{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-study-on-overfitting-in-deep-reinforcement","title":"A Study on Overfitting in Deep Reinforcement Learning","arxiv_id":"1804.06893","date":"2018-04-18","proceeding":null,"authors":["Chiyuan Zhang","Oriol Vinyals","Remi Munos","Samy Bengio"],"abstract":"Recent years have witnessed significant progresses in deep Reinforcement\nLearning (RL). Empowered with large scale neural networks, carefully designed\narchitectures, novel training algorithms and massively parallel computing\ndevices, researchers are able to attack many challenging RL problems. However,\nin machine learning, more training power comes with a potential risk of more\noverfitting. As deep RL techniques are being applied to critical problems such\nas healthcare and finance, it is important to understand the generalization\nbehaviors of the trained agents. In this paper, we conduct a systematic study\nof standard RL agents and find that they could overfit in various ways.\nMoreover, overfitting could happen \"robustly\": commonly used techniques in RL\nthat add stochasticity do not necessarily prevent or detect overfitting. In\nparticular, the same agents and learning algorithms could have drastically\ndifferent test performance, even when all of them achieve optimal rewards\nduring training. The observations call for more principled and careful\nevaluation protocols in RL. We conclude with a general discussion on\noverfitting in RL and a study of the generalization behaviors from the\nperspective of inductive bias.","url_abs":"http://arxiv.org/abs/1804.06893v2","url_pdf":"http://arxiv.org/pdf/1804.06893v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-study-on-overfitting-in-deep-reinforcement","repo_url":"https://github.com/oliviawl/image_classification_utkface","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"inductive-bias","task_name":"Inductive Bias"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.06893","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}