{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/leveraging-demonstrations-for-deep","title":"Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards","arxiv_id":"1707.08817","date":"2017-07-27","proceeding":null,"authors":["Mel Vecerik","Todd Hester","Jonathan Scholz","Fumin Wang","Olivier Pietquin","Bilal Piot","Nicolas Heess","Thomas Rothörl","Thomas Lampe","Martin Riedmiller"],"abstract":"We propose a general and model-free approach for Reinforcement Learning (RL)\non real robotics with sparse rewards. We build upon the Deep Deterministic\nPolicy Gradient (DDPG) algorithm to use demonstrations. Both demonstrations and\nactual interactions are used to fill a replay buffer and the sampling ratio\nbetween demonstrations and transitions is automatically tuned via a prioritized\nreplay mechanism. Typically, carefully engineered shaping rewards are required\nto enable the agents to efficiently explore on high dimensional control\nproblems such as robotics. They are also required for model-based acceleration\nmethods relying on local solvers such as iLQG (e.g. Guided Policy Search and\nNormalized Advantage Function). The demonstrations replace the need for\ncarefully engineered rewards, and reduce the exploration problem encountered by\nclassical RL approaches in these domains. Demonstrations are collected by a\nrobot kinesthetically force-controlled by a human demonstrator. Results on four\nsimulated insertion tasks show that DDPG from demonstrations out-performs DDPG,\nand does not require engineered rewards. Finally, we demonstrate the method on\na real robotics task consisting of inserting a clip (flexible object) into a\nrigid object.","url_abs":"http://arxiv.org/abs/1707.08817v2","url_pdf":"http://arxiv.org/pdf/1707.08817v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"leveraging-demonstrations-for-deep","repo_url":"https://github.com/MrSyee/pg-is-all-you-need","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"leveraging-demonstrations-for-deep","repo_url":"https://github.com/Zartris/TD3_continuous_control","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"leveraging-demonstrations-for-deep","repo_url":"https://github.com/Zartris/TD3_multi_agent_tennis","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"leveraging-demonstrations-for-deep","repo_url":"https://github.com/kairproject/kair_algorithms_draft","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"adam","method_name":"Adam"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"ddpg","method_name":"DDPG"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"experience-replay","method_name":"Experience Replay"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"weight-decay","method_name":"Weight Decay"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1707.08817","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}