{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/toybox-better-atari-environments-for-testing","title":"ToyBox: Better Atari Environments for Testing Reinforcement Learning Agents","arxiv_id":"1812.02850","date":"2018-12-06","proceeding":null,"authors":["John Foley","Emma Tosch","Kaleigh Clary","David Jensen"],"abstract":"It is a widely accepted principle that software without tests has bugs.\nTesting reinforcement learning agents is especially difficult because of the\nstochastic nature of both agents and environments, the complexity of\nstate-of-the-art models, and the sequential nature of their predictions.\nRecently, the Arcade Learning Environment (ALE) has become one of the most\nwidely used benchmark suites for deep learning research, and state-of-the-art\nReinforcement Learning (RL) agents have been shown to routinely equal or exceed\nhuman performance on many ALE tasks. Since ALE is based on emulation of\noriginal Atari games, the environment does not provide semantically meaningful\nrepresentations of internal game state. This means that ALE has limited utility\nas an environment for supporting testing or model introspection. We propose\nToyBox, a collection of reimplementations of these games that solves this\ncritical problem and enables robust testing of RL agents.","url_abs":"http://arxiv.org/abs/1812.02850v3","url_pdf":"http://arxiv.org/pdf/1812.02850v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"toybox-better-atari-environments-for-testing","repo_url":"https://github.com/KDL-umass/saliency_maps","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"toybox-better-atari-environments-for-testing","repo_url":"https://github.com/toybox-rs/Toybox","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"toybox-better-atari-environments-for-testing","repo_url":"https://github.com/toybox-rs/openai-baselines-envs","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"atari-games","task_name":"Atari Games"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.02850","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}