{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ai-safety-gridworlds","title":"AI Safety Gridworlds","arxiv_id":"1711.09883","date":"2017-11-27","proceeding":null,"authors":["Jan Leike","Miljan Martic","Victoria Krakovna","Pedro A. Ortega","Tom Everitt","Andrew Lefrancq","Laurent Orseau","Shane Legg"],"abstract":"We present a suite of reinforcement learning environments illustrating\nvarious safety properties of intelligent agents. These problems include safe\ninterruptibility, avoiding side effects, absent supervisor, reward gaming, safe\nexploration, as well as robustness to self-modification, distributional shift,\nand adversaries. To measure compliance with the intended safe behavior, we\nequip each environment with a performance function that is hidden from the\nagent. This allows us to categorize AI safety problems into robustness and\nspecification problems, depending on whether the performance function\ncorresponds to the observed reward function. We evaluate A2C and Rainbow, two\nrecent deep reinforcement learning agents, on our environments and show that\nthey are not able to solve them satisfactorily.","url_abs":"http://arxiv.org/abs/1711.09883v2","url_pdf":"http://arxiv.org/pdf/1711.09883v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ai-safety-gridworlds","repo_url":"https://github.com/deepmind/ai-safety-gridworlds","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"ai-safety-gridworlds","repo_url":"https://github.com/sharisun18/Absent_Supervisor_Env","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"safe-exploration","task_name":"Safe Exploration"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"a2c","method_name":"A2C"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1711.09883","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}