{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/high-dimensional-continuous-control-using","title":"High-Dimensional Continuous Control Using Generalized Advantage Estimation","arxiv_id":"1506.02438","date":"2015-06-08","proceeding":null,"authors":["John Schulman","Philipp Moritz","Sergey Levine","Michael Jordan","Pieter Abbeel"],"abstract":"Policy gradient methods are an appealing approach in reinforcement learning\nbecause they directly optimize the cumulative reward and can straightforwardly\nbe used with nonlinear function approximators such as neural networks. The two\nmain challenges are the large number of samples typically required, and the\ndifficulty of obtaining stable and steady improvement despite the\nnonstationarity of the incoming data. We address the first challenge by using\nvalue functions to substantially reduce the variance of policy gradient\nestimates at the cost of some bias, with an exponentially-weighted estimator of\nthe advantage function that is analogous to TD(lambda). We address the second\nchallenge by using trust region optimization procedure for both the policy and\nthe value function, which are represented by neural networks.\n  Our approach yields strong empirical results on highly challenging 3D\nlocomotion tasks, learning running gaits for bipedal and quadrupedal simulated\nrobots, and learning a policy for getting the biped to stand up from starting\nout lying on the ground. In contrast to a body of prior work that uses\nhand-crafted policy representations, our neural network policies map directly\nfrom raw kinematics to joint torques. Our algorithm is fully model-free, and\nthe amount of simulated experience required for the learning tasks on 3D bipeds\ncorresponds to 1-2 weeks of real time.","url_abs":"http://arxiv.org/abs/1506.02438v6","url_pdf":"http://arxiv.org/pdf/1506.02438v6.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/JonasRSV/PPO","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/UesugiErii/tf2-PPO-atari","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/adik993/ppo-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/alecfilios/Training-Intelligent-game-Agents-through-Competitive-Reinforcement-Learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/bentrevett/pytorch-rl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/dsinghnegi/atari_RL_agent","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/guillaumeboniface/reacher","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/magnusja/ppo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/mightypirate1/DRL-Tetris","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/morikatron/PPO","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/pat-coady/trpo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/ppocma/ppocma","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/shreyesss/PPO-implementation-keras-tensorflow","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/smj007/Breakout_A3C","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/taku-y/20181125-pybullet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/DLR-RM/stable-baselines3","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"high-dimensional-continuous-control-using","repo_url":"https://github.com/labmlai/annotated_deep_learning_paper_implementations","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"policy-gradient-methods","task_name":"Policy Gradient Methods"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"},{"task_slug":"continuous-control","task_name":"continuous-control"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1506.02438","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1506.02438"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/DLR-RM/stable-baselines3","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/morikatron/PPO","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/adik993/ppo-pytorch","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/alecfilios/Training-Intelligent-game-Agents-through-Competitive-Reinforcement-Learning","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/magnusja/ppo","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/smj007/Breakout_A3C","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ppocma/ppocma","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/shreyesss/PPO-implementation-keras-tensorflow","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/dsinghnegi/atari_RL_agent","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/UesugiErii/tf2-PPO-atari","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/labmlai/annotated_deep_learning_paper_implementations","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/JonasRSV/PPO","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/bentrevett/pytorch-rl","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/guillaumeboniface/reacher","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/pat-coady/trpo","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/taku-y/20181125-pybullet","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mightypirate1/DRL-Tetris","reach":{"status":"ok","spdx":"GPL-3.0"}}],"summary":{"ran_violates":1},"by_repo_kind":{},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"7e13cf9f394dd7cc","entry":"action_modifier","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"7e13cf9f394dd7cc"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}