{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/reward-certification-for-policy-smoothed","title":"Reward Certification for Policy Smoothed Reinforcement Learning","arxiv_id":"2312.06436","date":"2023-12-11","proceeding":null,"authors":["Ronghui Mu","Leandro Soriano Marcolino","Tianle Zhang","Yanghao Zhang","Xiaowei Huang","Wenjie Ruan"],"abstract":"Reinforcement Learning (RL) has achieved remarkable success in safety-critical areas, but it can be weakened by adversarial attacks. Recent studies have introduced \"smoothed policies\" in order to enhance its robustness. Yet, it is still challenging to establish a provable guarantee to certify the bound of its total reward. Prior methods relied primarily on computing bounds using Lipschitz continuity or calculating the probability of cumulative reward above specific thresholds. However, these techniques are only suited for continuous perturbations on the RL agent's observations and are restricted to perturbations bounded by the $l_2$-norm. To address these limitations, this paper proposes a general black-box certification method capable of directly certifying the cumulative reward of the smoothed policy under various $l_p$-norm bounded perturbations. Furthermore, we extend our methodology to certify perturbations on action spaces. Our approach leverages f-divergence to measure the distinction between the original distribution and the perturbed distribution, subsequently determining the certification bound by solving a convex optimisation problem. We provide a comprehensive theoretical analysis and run sufficient experiments in multiple environments. Our results show that our method not only improves the certified lower bound of mean cumulative reward but also demonstrates better efficiency than state-of-the-art techniques.","url_abs":"https://arxiv.org/abs/2312.06436v2","url_pdf":"https://arxiv.org/pdf/2312.06436v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"reward-certification-for-policy-smoothed","repo_url":"https://github.com/trustai/receps","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2312.06436","atlas_url":"https://app.syntology.ai/?focus=2312.06436","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2312.06436"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/trustai/receps","reach":{"status":"ok"}}],"summary":{"ran":5,"unverified":6},"by_repo_kind":{"official":{"samples":11,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":11,"samples":[{"code_sha256_prefix":"38ff021176e23ea8","entry":"action_attack","repo":"trustai/receps","repo_kind":"official","path":"Freeway/action_attack_free.py","file_url":"https://github.com/trustai/receps/blob/HEAD/Freeway/action_attack_free.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"38ff021176e23ea8"}},{"code_sha256_prefix":"0effa27eababf646","entry":"choose_action","repo":"trustai/receps","repo_kind":"official","path":"Cartpole/cart-train-action-smoot.py","file_url":"https://github.com/trustai/receps/blob/HEAD/Cartpole/cart-train-action-smoot.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0effa27eababf646"}},{"code_sha256_prefix":"2d84acacb6f9890e","entry":"get_cvar_cert_time_t","repo":"trustai/receps","repo_kind":"official","path":"Cartpole/cartpole_multiframe_plot_attacks.py","file_url":"https://github.com/trustai/receps/blob/HEAD/Cartpole/cartpole_multiframe_plot_attacks.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2d84acacb6f9890e"}},{"code_sha256_prefix":"ec12fc9d760fc3a6","entry":"soft_smooth_fun","repo":"trustai/receps","repo_kind":"official","path":"Cartpole/cart_attackl1.py","file_url":"https://github.com/trustai/receps/blob/HEAD/Cartpole/cart_attackl1.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ec12fc9d760fc3a6"}},{"code_sha256_prefix":"a3ac61a71353c977","entry":"threshold","repo":"trustai/receps","repo_kind":"official","path":"Freeway/action_attack_free.py","file_url":"https://github.com/trustai/receps/blob/HEAD/Freeway/action_attack_free.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a3ac61a71353c977"}},{"code_sha256_prefix":"9fe7c510e1252f0b","entry":"action_attack","repo":"trustai/receps","repo_kind":"official","path":"Cartpole/action_attack_cartpole.py","file_url":"https://github.com/trustai/receps/blob/HEAD/Cartpole/action_attack_cartpole.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9fe7c510e1252f0b"}},{"code_sha256_prefix":"b2e14f39728a999d","entry":"attack","repo":"trustai/receps","repo_kind":"official","path":"Cartpole/cart_attackl1.py","file_url":"https://github.com/trustai/receps/blob/HEAD/Cartpole/cart_attackl1.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b2e14f39728a999d"}},{"code_sha256_prefix":"ff6cfb13e671861b","entry":"attack","repo":"trustai/receps","repo_kind":"official","path":"Cartpole/cartpole_multiframe_attack_smooth.py","file_url":"https://github.com/trustai/receps/blob/HEAD/Cartpole/cartpole_multiframe_attack_smooth.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ff6cfb13e671861b"}},{"code_sha256_prefix":"dfb7ee9573f8543d","entry":"attack_clean","repo":"trustai/receps","repo_kind":"official","path":"Cartpole/cartpole_multiframe_attack.py","file_url":"https://github.com/trustai/receps/blob/HEAD/Cartpole/cartpole_multiframe_attack.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"dfb7ee9573f8543d"}},{"code_sha256_prefix":"c8b6dadbdaac9a3f","entry":"attack_clean","repo":"trustai/receps","repo_kind":"official","path":"Cartpole/cartpole_simple_attack.py","file_url":"https://github.com/trustai/receps/blob/HEAD/Cartpole/cartpole_simple_attack.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"c8b6dadbdaac9a3f"}},{"code_sha256_prefix":"81d7081fdee68ac2","entry":"threshold","repo":"trustai/receps","repo_kind":"official","path":"Cartpole/action_attack_cartpole.py","file_url":"https://github.com/trustai/receps/blob/HEAD/Cartpole/action_attack_cartpole.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"81d7081fdee68ac2"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}