{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-run-challenge-solutions-adapting","title":"Learning to Run challenge solutions: Adapting reinforcement learning methods for neuromusculoskeletal environments","arxiv_id":"1804.00361","date":"2018-04-02","proceeding":null,"authors":["Łukasz Kidziński","Sharada Prasanna Mohanty","Carmichael Ong","Zhewei Huang","Shuchang Zhou","Anton Pechenko","Adam Stelmaszczyk","Piotr Jarosik","Mikhail Pavlov","Sergey Kolesnikov","Sergey Plis","Zhibo Chen","Zhizheng Zhang","Jiale Chen","Jun Shi","Zhuobin Zheng","Chun Yuan","Zhihui Lin","Henryk Michalewski","Piotr Miłoś","Błażej Osiński","Andrew Melnik","Malte Schilling","Helge Ritter","Sean Carroll","Jennifer Hicks","Sergey Levine","Marcel Salathé","Scott Delp"],"abstract":"In the NIPS 2017 Learning to Run challenge, participants were tasked with\nbuilding a controller for a musculoskeletal model to make it run as fast as\npossible through an obstacle course. Top participants were invited to describe\ntheir algorithms. In this work, we present eight solutions that used deep\nreinforcement learning approaches, based on algorithms such as Deep\nDeterministic Policy Gradient, Proximal Policy Optimization, and Trust Region\nPolicy Optimization. Many solutions use similar relaxations and heuristics,\nsuch as reward shaping, frame skipping, discretization of the action space,\nsymmetry, and policy blending. However, each of the eight teams implemented\ndifferent modifications of the known algorithms.","url_abs":"http://arxiv.org/abs/1804.00361v1","url_pdf":"http://arxiv.org/pdf/1804.00361v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-run-challenge-solutions-adapting","repo_url":"https://github.com/AdamStelmaszczyk/learning2run","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"learning-to-run-challenge-solutions-adapting","repo_url":"https://github.com/megvii-research/NIPS2017-LearningToRunACE","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.00361","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1804.00361"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/megvii-research/NIPS2017-LearningToRunACE","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/AdamStelmaszczyk/learning2run","reach":null}],"summary":{"ran_honours":1,"unverified":2},"by_repo_kind":{"official":{"samples":3,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"387ed23192d3abe4","entry":"obstacle_x_importance","repo":"AdamStelmaszczyk/learning2run","repo_kind":"official","path":"baselines/baselines/pposgd/pposgd_simple.py","file_url":"https://github.com/AdamStelmaszczyk/learning2run/blob/HEAD/baselines/baselines/pposgd/pposgd_simple.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"387ed23192d3abe4"}},{"code_sha256_prefix":"a734b6d74ae254b4","entry":"add_accelerations","repo":"AdamStelmaszczyk/learning2run","repo_kind":"official","path":"baselines/baselines/pposgd/pposgd_simple.py","file_url":"https://github.com/AdamStelmaszczyk/learning2run/blob/HEAD/baselines/baselines/pposgd/pposgd_simple.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a734b6d74ae254b4"}},{"code_sha256_prefix":"4c46daa0094a56d0","entry":"add_velocities","repo":"AdamStelmaszczyk/learning2run","repo_kind":"official","path":"baselines/baselines/pposgd/pposgd_simple.py","file_url":"https://github.com/AdamStelmaszczyk/learning2run/blob/HEAD/baselines/baselines/pposgd/pposgd_simple.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4c46daa0094a56d0"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}