{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/open-rl-benchmark-comprehensive-tracked","title":"Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning","arxiv_id":"2402.03046","date":"2024-02-05","proceeding":null,"authors":["Shengyi Huang","Quentin Gallouédec","Florian Felten","Antonin Raffin","Rousslan Fernand Julien Dossa","Yanxiao Zhao","Ryan Sullivan","Viktor Makoviychuk","Denys Makoviichuk","Mohamad H. Danesh","Cyril Roumégous","Jiayi Weng","Chufan Chen","Md Masudur Rahman","João G. M. Araújo","Guorui Quan","Daniel Tan","Timo Klein","Rujikorn Charakorn","Mark Towers","Yann Berthelot","Kinal Mehta","Dipam Chakraborty","Arjun KG","Valentin Charraut","Chang Ye","Zichen Liu","Lucas N. Alegre","Alexander Nikulin","Xiao Hu","Tianlin Liu","Jongwook Choi","Brent Yi"],"abstract":"In many Reinforcement Learning (RL) papers, learning curves are useful indicators to measure the effectiveness of RL algorithms. However, the complete raw data of the learning curves are rarely available. As a result, it is usually necessary to reproduce the experiments from scratch, which can be time-consuming and error-prone. We present Open RL Benchmark, a set of fully tracked RL experiments, including not only the usual data such as episodic return, but also all algorithm-specific and system metrics. Open RL Benchmark is community-driven: anyone can download, use, and contribute to the data. At the time of writing, more than 25,000 runs have been tracked, for a cumulative duration of more than 8 years. Open RL Benchmark covers a wide range of RL libraries and reference implementations. Special care is taken to ensure that each experiment is precisely reproducible by providing not only the full parameters, but also the versions of the dependencies used to generate it. In addition, Open RL Benchmark comes with a command-line interface (CLI) for easy fetching and generating figures to present the results. In this document, we include two case studies to demonstrate the usefulness of Open RL Benchmark in practice. To the best of our knowledge, Open RL Benchmark is the first RL benchmark of its kind, and the authors hope that it will improve and facilitate the work of researchers in the field.","url_abs":"https://arxiv.org/abs/2402.03046v1","url_pdf":"https://arxiv.org/pdf/2402.03046v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"open-rl-benchmark-comprehensive-tracked","repo_url":"https://github.com/openrlbenchmark/openrlbenchmark","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"jax","reach":null}],"tasks":[{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2402.03046","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2402.03046"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/openrlbenchmark/openrlbenchmark","reach":null}],"summary":{"ran_violates":1},"by_repo_kind":{"listed":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"6972dfe7485c9556","entry":"convert","repo":"openrlbenchmark/openrlbenchmark","repo_kind":"listed","path":"openrlbenchmark/rlops.py","file_url":"https://github.com/openrlbenchmark/openrlbenchmark/blob/HEAD/openrlbenchmark/rlops.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6972dfe7485c9556"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}