{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-robust-options-by-conditional-value","title":"Learning Robust Options by Conditional Value at Risk Optimization","arxiv_id":"1905.09191","date":"2019-05-22","proceeding":"NeurIPS 2019 12","authors":["Takuya Hiraoka","Takahisa Imagawa","Tatsuya Mori","Takashi Onishi","Yoshimasa Tsuruoka"],"abstract":"Options are generally learned by using an inaccurate environment model (or simulator), which contains uncertain model parameters. While there are several methods to learn options that are robust against the uncertainty of model parameters, these methods only consider either the worst case or the average (ordinary) case for learning options. This limited consideration of the cases often produces options that do not work well in the unconsidered case. In this paper, we propose a conditional value at risk (CVaR)-based method to learn options that work well in both the average and worst cases. We extend the CVaR-based policy gradient method proposed by Chow and Ghavamzadeh (2014) to deal with robust Markov decision processes and then apply the extended method to learning robust options. We conduct experiments to evaluate our method in multi-joint robot control tasks (HopperIceBlock, Half-Cheetah, and Walker2D). Experimental results show that our method produces options that 1) give better worst-case performance than the options learned only to minimize the average-case loss, and 2) give better average-case performance than the options learned only to minimize the worst-case loss.","url_abs":"https://arxiv.org/abs/1905.09191v4","url_pdf":"https://arxiv.org/pdf/1905.09191v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-robust-options-by-conditional-value","repo_url":"https://github.com/TakuyaHiraoka/Learning-Robust-Options-by-Conditional-Value-at-Risk-Optimization","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1905.09191","atlas_url":"https://app.syntology.ai/?focus=1905.09191","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1905.09191"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/TakuyaHiraoka/Learning-Robust-Options-by-Conditional-Value-at-Risk-Optimization","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_honours":1,"unverified":3},"by_repo_kind":{"official":{"samples":4,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"0173e4280d9a9884","entry":"flatten_lists","repo":"TakuyaHiraoka/Learning-Robust-Options-by-Conditional-Value-at-Risk-Optimization","repo_kind":"official","path":"LearningMoreRobustOption/pposgd_simple.py","file_url":"https://github.com/TakuyaHiraoka/Learning-Robust-Options-by-Conditional-Value-at-Risk-Optimization/blob/HEAD/LearningMoreRobustOption/pposgd_simple.py","link_basis":"harvester_set","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0173e4280d9a9884"}},{"code_sha256_prefix":"c87dbab117f0acc8","entry":"dense3D2","repo":"TakuyaHiraoka/Learning-Robust-Options-by-Conditional-Value-at-Risk-Optimization","repo_kind":"official","path":"LearningMoreRobustOption/mlp_policy.py","file_url":"https://github.com/TakuyaHiraoka/Learning-Robust-Options-by-Conditional-Value-at-Risk-Optimization/blob/HEAD/LearningMoreRobustOption/mlp_policy.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c87dbab117f0acc8"}},{"code_sha256_prefix":"9ba2f70546cf0616","entry":"find_all_key_files_path","repo":"TakuyaHiraoka/Learning-Robust-Options-by-Conditional-Value-at-Risk-Optimization","repo_kind":"official","path":"Scores4EachParameterPerturbation.py","file_url":"https://github.com/TakuyaHiraoka/Learning-Robust-Options-by-Conditional-Value-at-Risk-Optimization/blob/HEAD/Scores4EachParameterPerturbation.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9ba2f70546cf0616"}},{"code_sha256_prefix":"139b7d9b5bb0e203","entry":"find_all_key_files_path","repo":"TakuyaHiraoka/Learning-Robust-Options-by-Conditional-Value-at-Risk-Optimization","repo_kind":"official","path":"EvalAverageCVaR.py","file_url":"https://github.com/TakuyaHiraoka/Learning-Robust-Options-by-Conditional-Value-at-Risk-Optimization/blob/HEAD/EvalAverageCVaR.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"139b7d9b5bb0e203"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}