{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/benchmarking-reinforcement-learning","title":"Benchmarking Reinforcement Learning Algorithms on Real-World Robots","arxiv_id":"1809.07731","date":"2018-09-20","proceeding":null,"authors":["A. Rupam Mahmood","Dmytro Korenkevych","Gautham Vasan","William Ma","James Bergstra"],"abstract":"Through many recent successes in simulation, model-free reinforcement\nlearning has emerged as a promising approach to solving continuous control\nrobotic tasks. The research community is now able to reproduce, analyze and\nbuild quickly on these results due to open source implementations of learning\nalgorithms and simulated benchmark tasks. To carry forward these successes to\nreal-world applications, it is crucial to withhold utilizing the unique\nadvantages of simulations that do not transfer to the real world and experiment\ndirectly with physical robots. However, reinforcement learning research with\nphysical robots faces substantial resistance due to the lack of benchmark tasks\nand supporting source code. In this work, we introduce several reinforcement\nlearning tasks with multiple commercially available robots that present varying\nlevels of learning difficulty, setup, and repeatability. On these tasks, we\ntest the learning performance of off-the-shelf implementations of four\nreinforcement learning algorithms and analyze sensitivity to their\nhyper-parameters to determine their readiness for applications in various\nreal-world tasks. Our results show that with a careful setup of the task\ninterface and computations, some of these implementations can be readily\napplicable to physical robots. We find that state-of-the-art learning\nalgorithms are highly sensitive to their hyper-parameters and their relative\nordering does not transfer across tasks, indicating the necessity of re-tuning\nthem for each task for best performance. On the other hand, the best\nhyper-parameter configuration from one task may often result in effective\nlearning on held-out tasks even with different robots, providing a reasonable\ndefault. We make the benchmark tasks publicly available to enhance\nreproducibility in real-world reinforcement learning.","url_abs":"http://arxiv.org/abs/1809.07731v1","url_pdf":"http://arxiv.org/pdf/1809.07731v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"benchmarking-reinforcement-learning","repo_url":"https://github.com/kindredresearch/SenseAct","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"paper_slug":"benchmarking-reinforcement-learning","repo_url":"https://github.com/dti-research/SenseAct","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1809.07731","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1809.07731"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/kindredresearch/SenseAct","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/dti-research/SenseAct","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"summary":{"unverified":2},"by_repo_kind":{"official":{"samples":2,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"7fd22c31ecb07bdd","entry":"get_random_state_array","repo":"kindredresearch/SenseAct","repo_kind":"official","path":"senseact/utils.py","file_url":"https://github.com/kindredresearch/SenseAct/blob/HEAD/senseact/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"7fd22c31ecb07bdd"}},{"code_sha256_prefix":"007769cae2550336","entry":"get_random_state_from_array","repo":"kindredresearch/SenseAct","repo_kind":"official","path":"senseact/utils.py","file_url":"https://github.com/kindredresearch/SenseAct/blob/HEAD/senseact/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"007769cae2550336"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}