{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/benchmarking-offline-reinforcement-learning-1","title":"Benchmarking Offline Reinforcement Learning on Real-Robot Hardware","arxiv_id":"2307.15690","date":"2023-07-28","proceeding":null,"authors":["Nico Gürtler","Sebastian Blaes","Pavel Kolev","Felix Widmaier","Manuel Wüthrich","Stefan Bauer","Bernhard Schölkopf","Georg Martius"],"abstract":"Learning policies from previously recorded data is a promising direction for real-world robotics tasks, as online learning is often infeasible. Dexterous manipulation in particular remains an open problem in its general form. The combination of offline reinforcement learning with large diverse datasets, however, has the potential to lead to a breakthrough in this challenging domain analogously to the rapid progress made in supervised learning in recent years. To coordinate the efforts of the research community toward tackling this problem, we propose a benchmark including: i) a large collection of data for offline learning from a dexterous manipulation platform on two tasks, obtained with capable RL agents trained in simulation; ii) the option to execute learned policies on a real-world robotic system and a simulation for efficient debugging. We evaluate prominent open-sourced offline reinforcement learning algorithms on the datasets and provide a reproducible experimental setup for offline reinforcement learning on real systems.","url_abs":"https://arxiv.org/abs/2307.15690v1","url_pdf":"https://arxiv.org/pdf/2307.15690v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"benchmarking-offline-reinforcement-learning-1","repo_url":"https://github.com/rr-learning/trifinger-rl-example","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"paper_slug":"benchmarking-offline-reinforcement-learning-1","repo_url":"https://github.com/rr-learning/trifinger_rl_datasets","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2307.15690","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2307.15690"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/rr-learning/trifinger_rl_datasets","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rr-learning/trifinger-rl-example","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"summary":{"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"f227eb1a410a4522","entry":"get_keypoints_from_pose","repo":"rr-learning/trifinger_rl_datasets","repo_kind":"official","path":"trifinger_rl_datasets/utils.py","file_url":"https://github.com/rr-learning/trifinger_rl_datasets/blob/HEAD/trifinger_rl_datasets/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"f227eb1a410a4522"}},{"code_sha256_prefix":"d6becba6a8b4f792","entry":"to_quat","repo":"rr-learning/trifinger_rl_datasets","repo_kind":"official","path":"trifinger_rl_datasets/utils.py","file_url":"https://github.com/rr-learning/trifinger_rl_datasets/blob/HEAD/trifinger_rl_datasets/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"d6becba6a8b4f792"}},{"code_sha256_prefix":"850012dbaa8b0fff","entry":"to_world_space","repo":"rr-learning/trifinger_rl_datasets","repo_kind":"official","path":"trifinger_rl_datasets/utils.py","file_url":"https://github.com/rr-learning/trifinger_rl_datasets/blob/HEAD/trifinger_rl_datasets/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"850012dbaa8b0fff"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}