{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/action-robust-reinforcement-learning-and","title":"Action Robust Reinforcement Learning and Applications in Continuous Control","arxiv_id":"1901.09184","date":"2019-01-26","proceeding":null,"authors":["Chen Tessler","Yonathan Efroni","Shie Mannor"],"abstract":"A policy is said to be robust if it maximizes the reward while considering a bad, or even adversarial, model. In this work we formalize two new criteria of robustness to action uncertainty. Specifically, we consider two scenarios in which the agent attempts to perform an action $a$, and (i) with probability $\\alpha$, an alternative adversarial action $\\bar a$ is taken, or (ii) an adversary adds a perturbation to the selected action in the case of continuous action space. We show that our criteria are related to common forms of uncertainty in robotics domains, such as the occurrence of abrupt forces, and suggest algorithms in the tabular case. Building on the suggested algorithms, we generalize our approach to deep reinforcement learning (DRL) and provide extensive experiments in the various MuJoCo domains. Our experiments show that not only does our approach produce robust policies, but it also improves the performance in the absence of perturbations. This generalization indicates that action-robustness can be thought of as implicit regularization in RL problems.","url_abs":"https://arxiv.org/abs/1901.09184v2","url_pdf":"https://arxiv.org/pdf/1901.09184v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"action-robust-reinforcement-learning-and","repo_url":"https://github.com/tesslerc/ActionRobustRL","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"action-robust-reinforcement-learning-and","repo_url":"https://github.com/icml2019-anonymous-author/Action-Robust-Reinforcement-Learning","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"mujoco","task_name":"MuJoCo"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1901.09184","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1901.09184"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/icml2019-anonymous-author/Action-Robust-Reinforcement-Learning","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/tesslerc/ActionRobustRL","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_honours":1,"unverified":5},"by_repo_kind":{"official":{"samples":6,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"3ebb420a5d10cc64","entry":"update_mean_var_count_from_moments","repo":"icml2019-anonymous-author/Action-Robust-Reinforcement-Learning","repo_kind":"official","path":"utils.py","file_url":"https://github.com/icml2019-anonymous-author/Action-Robust-Reinforcement-Learning/blob/HEAD/utils.py","link_basis":"harvester_set","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3ebb420a5d10cc64"}},{"code_sha256_prefix":"c088fbf627df2a54","entry":"ddpg_distance_metric","repo":"icml2019-anonymous-author/Action-Robust-Reinforcement-Learning","repo_kind":"official","path":"param_noise.py","file_url":"https://github.com/icml2019-anonymous-author/Action-Robust-Reinforcement-Learning/blob/HEAD/param_noise.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c088fbf627df2a54"}},{"code_sha256_prefix":"00277beee4c6e0a9","entry":"denormalize","repo":"icml2019-anonymous-author/Action-Robust-Reinforcement-Learning","repo_kind":"official","path":"ddpg.py","file_url":"https://github.com/icml2019-anonymous-author/Action-Robust-Reinforcement-Learning/blob/HEAD/ddpg.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"00277beee4c6e0a9"}},{"code_sha256_prefix":"6fd5995ea3f14938","entry":"normalize","repo":"icml2019-anonymous-author/Action-Robust-Reinforcement-Learning","repo_kind":"official","path":"ddpg.py","file_url":"https://github.com/icml2019-anonymous-author/Action-Robust-Reinforcement-Learning/blob/HEAD/ddpg.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6fd5995ea3f14938"}},{"code_sha256_prefix":"cd302ca126035e5c","entry":"normalize","repo":"icml2019-anonymous-author/Action-Robust-Reinforcement-Learning","repo_kind":"official","path":"normalized_actions.py","file_url":"https://github.com/icml2019-anonymous-author/Action-Robust-Reinforcement-Learning/blob/HEAD/normalized_actions.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"cd302ca126035e5c"}},{"code_sha256_prefix":"bc3dde5f5f1ac8e2","entry":"update_mean_var_count_from_moments","repo":"icml2019-anonymous-author/Action-Robust-Reinforcement-Learning","repo_kind":"official","path":"normalized_actions.py","file_url":"https://github.com/icml2019-anonymous-author/Action-Robust-Reinforcement-Learning/blob/HEAD/normalized_actions.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"bc3dde5f5f1ac8e2"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}