{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/policy-transfer-with-strategy-optimization","title":"Policy Transfer with Strategy Optimization","arxiv_id":"1810.05751","date":"2018-10-12","proceeding":"ICLR 2019 5","authors":["Wenhao Yu","C. Karen Liu","Greg Turk"],"abstract":"Computer simulation provides an automatic and safe way for training robotic\ncontrol policies to achieve complex tasks such as locomotion. However, a policy\ntrained in simulation usually does not transfer directly to the real hardware\ndue to the differences between the two environments. Transfer learning using\ndomain randomization is a promising approach, but it usually assumes that the\ntarget environment is close to the distribution of the training environments,\nthus relying heavily on accurate system identification. In this paper, we\npresent a different approach that leverages domain randomization for\ntransferring control policies to unknown environments. The key idea that,\ninstead of learning a single policy in the simulation, we simultaneously learn\na family of policies that exhibit different behaviors. When tested in the\ntarget environment, we directly search for the best policy in the family based\non the task performance, without the need to identify the dynamic parameters.\nWe evaluate our method on five simulated robotic control problems with\ndifferent discrepancies in the training and testing environment and demonstrate\nthat our method can overcome larger modeling errors compared to training a\nrobust policy or an adaptive policy.","url_abs":"http://arxiv.org/abs/1810.05751v2","url_pdf":"http://arxiv.org/pdf/1810.05751v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"policy-transfer-with-strategy-optimization","repo_url":"https://github.com/vincentyu68/policy_transfer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.05751","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}