{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-to-pick-the-domain-randomization","title":"How to pick the domain randomization parameters for sim-to-real transfer of reinforcement learning policies?","arxiv_id":"1903.11774","date":"2019-03-28","proceeding":null,"authors":["Quan Vuong","Sharad Vikram","Hao Su","Sicun Gao","Henrik I. Christensen"],"abstract":"Recently, reinforcement learning (RL) algorithms have demonstrated remarkable\nsuccess in learning complicated behaviors from minimally processed input.\nHowever, most of this success is limited to simulation. While there are\npromising successes in applying RL algorithms directly on real systems, their\nperformance on more complex systems remains bottle-necked by the relative data\ninefficiency of RL algorithms. Domain randomization is a promising direction of\nresearch that has demonstrated impressive results using RL algorithms to\ncontrol real robots. At a high level, domain randomization works by training a\npolicy on a distribution of environmental conditions in simulation. If the\nenvironments are diverse enough, then the policy trained on this distribution\nwill plausibly generalize to the real world. A human-specified design choice in\ndomain randomization is the form and parameters of the distribution of\nsimulated environments. It is unclear how to the best pick the form and\nparameters of this distribution and prior work uses hand-tuned distributions.\nThis extended abstract demonstrates that the choice of the distribution plays a\nmajor role in the performance of the trained policies in the real world and\nthat the parameter of this distribution can be optimized to maximize the\nperformance of the trained policies in the real world","url_abs":"http://arxiv.org/abs/1903.11774v1","url_pdf":"http://arxiv.org/pdf/1903.11774v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-to-pick-the-domain-randomization","repo_url":"https://github.com/quanvuong/domain_randomization","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.11774","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1903.11774"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/quanvuong/domain_randomization","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"1def8a0dfa93b973","entry":"make_mujoco_env","repo":"quanvuong/domain_randomization","repo_kind":"official","path":"dr/ppo/utils.py","file_url":"https://github.com/quanvuong/domain_randomization/blob/HEAD/dr/ppo/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1def8a0dfa93b973"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}