{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/run-skeleton-run-skeletal-model-in-a-physics","title":"Run, skeleton, run: skeletal model in a physics-based simulation","arxiv_id":"1711.06922","date":"2017-11-18","proceeding":null,"authors":["Mikhail Pavlov","Sergey Kolesnikov","Sergey M. Plis"],"abstract":"In this paper, we present our approach to solve a physics-based reinforcement\nlearning challenge \"Learning to Run\" with objective to train\nphysiologically-based human model to navigate a complex obstacle course as\nquickly as possible. The environment is computationally expensive, has a\nhigh-dimensional continuous action space and is stochastic. We benchmark state\nof the art policy-gradient methods and test several improvements, such as layer\nnormalization, parameter noise, action and state reflecting, to stabilize\ntraining and improve its sample-efficiency. We found that the Deep\nDeterministic Policy Gradient method is the most efficient method for this\nenvironment and the improvements we have introduced help to stabilize training.\nLearned models are able to generalize to new physical scenarios, e.g. different\nobstacle courses.","url_abs":"http://arxiv.org/abs/1711.06922v2","url_pdf":"http://arxiv.org/pdf/1711.06922v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"run-skeleton-run-skeletal-model-in-a-physics","repo_url":"https://github.com/Scitator/Run-Skeleton-Run","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"navigate","task_name":"Navigate"},{"task_slug":"policy-gradient-methods","task_name":"Policy Gradient Methods"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1711.06922","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}