{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/risk-aware-reward-shaping-of-reinforcement","title":"Risk-Aware Reward Shaping of Reinforcement Learning Agents for Autonomous Driving","arxiv_id":"2306.03220","date":"2023-06-05","proceeding":null,"authors":["Lin-Chi Wu","Zengjie Zhang","Sofie Haesaert","Zhiqiang Ma","Zhiyong Sun"],"abstract":"Reinforcement learning (RL) is an effective approach to motion planning in autonomous driving, where an optimal driving policy can be automatically learned using the interaction data with the environment. Nevertheless, the reward function for an RL agent, which is significant to its performance, is challenging to be determined. The conventional work mainly focuses on rewarding safe driving states but does not incorporate the awareness of risky driving behaviors of the vehicles. In this paper, we investigate how to use risk-aware reward shaping to leverage the training and test performance of RL agents in autonomous driving. Based on the essential requirements that prescribe the safety specifications for general autonomous driving in practice, we propose additional reshaped reward terms that encourage exploration and penalize risky driving behaviors. A simulation study in OpenAI Gym indicates the advantage of risk-aware reward shaping for various RL agents. Also, we point out that proximal policy optimization (PPO) is likely to be the best RL method that works with risk-aware reward shaping.","url_abs":"https://arxiv.org/abs/2306.03220v2","url_pdf":"https://arxiv.org/pdf/2306.03220v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"risk-aware-reward-shaping-of-reinforcement","repo_url":"https://github.com/zhang-zengjie/code_2023_iecon_shaping_wu","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"tasks":[{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"motion-planning","task_name":"Motion Planning"},{"task_slug":"openai-gym","task_name":"OpenAI Gym"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"test","method_name":"Test"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2306.03220","atlas_url":"https://app.syntology.ai/?focus=2306.03220","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2306.03220"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zhang-zengjie/code_2023_iecon_shaping_wu","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"summary":{"ran_honours":1,"unverified":3},"by_repo_kind":{"official":{"samples":4,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"24ee38e114b38bce","entry":"hidden_init","repo":"zhang-zengjie/code_2023_iecon_shaping_wu","repo_kind":"official","path":"agent/model_ddpg.py","file_url":"https://github.com/zhang-zengjie/code_2023_iecon_shaping_wu/blob/HEAD/agent/model_ddpg.py","link_basis":"plan_row","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"24ee38e114b38bce"}},{"code_sha256_prefix":"2c5c3ba91ff2aacd","entry":"get_unique_numbers","repo":"zhang-zengjie/code_2023_iecon_shaping_wu","repo_kind":"official","path":"utility.py","file_url":"https://github.com/zhang-zengjie/code_2023_iecon_shaping_wu/blob/HEAD/utility.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"2c5c3ba91ff2aacd"}},{"code_sha256_prefix":"d618b702fead2858","entry":"score_action_save","repo":"zhang-zengjie/code_2023_iecon_shaping_wu","repo_kind":"official","path":"utility.py","file_url":"https://github.com/zhang-zengjie/code_2023_iecon_shaping_wu/blob/HEAD/utility.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"d618b702fead2858"}},{"code_sha256_prefix":"faecb51a58932147","entry":"score_save","repo":"zhang-zengjie/code_2023_iecon_shaping_wu","repo_kind":"official","path":"utility.py","file_url":"https://github.com/zhang-zengjie/code_2023_iecon_shaping_wu/blob/HEAD/utility.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"faecb51a58932147"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}