{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bipedal-walking-robot-using-deep","title":"Bipedal Walking Robot using Deep Deterministic Policy Gradient","arxiv_id":"1807.05924","date":"2018-07-16","proceeding":null,"authors":["Arun Kumar","Navneet Paul","S. N. Omkar"],"abstract":"Machine learning algorithms have found several applications in the field of\nrobotics and control systems. The control systems community has started to show\ninterest towards several machine learning algorithms from the sub-domains such\nas supervised learning, imitation learning and reinforcement learning to\nachieve autonomous control and intelligent decision making. Amongst many\ncomplex control problems, stable bipedal walking has been the most challenging\nproblem. In this paper, we present an architecture to design and simulate a\nplanar bipedal walking robot(BWR) using a realistic robotics simulator, Gazebo.\nThe robot demonstrates successful walking behaviour by learning through several\nof its trial and errors, without any prior knowledge of itself or the world\ndynamics. The autonomous walking of the BWR is achieved using reinforcement\nlearning algorithm called Deep Deterministic Policy Gradient(DDPG). DDPG is one\nof the algorithms for learning controls in continuous action spaces. After\ntraining the model in simulation, it was observed that, with a proper shaped\nreward function, the robot achieved faster walking or even rendered a running\ngait with an average speed of 0.83 m/s. The gait pattern of the bipedal walker\nwas compared with the actual human walking pattern. The results show that the\nbipedal walking pattern had similar characteristics to that of a human walking\npattern. The video presenting our experiment is available at\nhttps://goo.gl/NHXKqR.","url_abs":"http://arxiv.org/abs/1807.05924v2","url_pdf":"http://arxiv.org/pdf/1807.05924v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bipedal-walking-robot-using-deep","repo_url":"https://github.com/nav74neet/ddpg4biped","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"bipedal-walking-robot-using-deep","repo_url":"https://github.com/nav74neet/ddpg_biped","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"bipedal-walking-robot-using-deep","repo_url":"https://github.com/nav74neet/rl4biped","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"imitation-learning","task_name":"Imitation Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"adam","method_name":"Adam"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"ddpg","method_name":"DDPG"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"experience-replay","method_name":"Experience Replay"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"speed","method_name":"SPEED"},{"method_slug":"weight-decay","method_name":"Weight Decay"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}