{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scalable-trust-region-method-for-deep","title":"Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation","arxiv_id":"1708.05144","date":"2017-08-17","proceeding":"NeurIPS 2017 12","authors":["Yuhuai Wu","Elman Mansimov","Shun Liao","Roger Grosse","Jimmy Ba"],"abstract":"In this work, we propose to apply trust region optimization to deep\nreinforcement learning using a recently proposed Kronecker-factored\napproximation to the curvature. We extend the framework of natural policy\ngradient and propose to optimize both the actor and the critic using\nKronecker-factored approximate curvature (K-FAC) with trust region; hence we\ncall our method Actor Critic using Kronecker-Factored Trust Region (ACKTR). To\nthe best of our knowledge, this is the first scalable trust region natural\ngradient method for actor-critic methods. It is also a method that learns\nnon-trivial tasks in continuous control as well as discrete control policies\ndirectly from raw pixel inputs. We tested our approach across discrete domains\nin Atari games as well as continuous domains in the MuJoCo environment. With\nthe proposed methods, we are able to achieve higher rewards and a 2- to 3-fold\nimprovement in sample efficiency on average, compared to previous\nstate-of-the-art on-policy actor-critic methods. Code is available at\nhttps://github.com/openai/baselines","url_abs":"http://arxiv.org/abs/1708.05144v2","url_pdf":"http://arxiv.org/pdf/1708.05144v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"scalable-trust-region-method-for-deep","repo_url":"https://github.com/openai/baselines","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null},{"paper_slug":"scalable-trust-region-method-for-deep","repo_url":"https://github.com/angelolovatto/raylab","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"scalable-trust-region-method-for-deep","repo_url":"https://github.com/ermongroup/higher_order_invariance","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"scalable-trust-region-method-for-deep","repo_url":"https://github.com/ghostFaceKillah/expert","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"scalable-trust-region-method-for-deep","repo_url":"https://github.com/jrobine/actor-critic","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"scalable-trust-region-method-for-deep","repo_url":"https://github.com/monte-ai/expert-augmented-acktr","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"scalable-trust-region-method-for-deep","repo_url":"https://github.com/saschaschramm/Pong","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"scalable-trust-region-method-for-deep","repo_url":"https://github.com/hill-a/stable-baselines","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"atari-games","task_name":"Atari Games"},{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"mujoco","task_name":"MuJoCo"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"actkr","method_name":"ACTKR"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"entropy-regularization","method_name":"Entropy Regularization"}],"datasets_introduced":[],"methods_introduced":[{"slug":"actkr","name":"ACTKR","full_name":"ACTKR"}],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1708.05144","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1708.05144"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/hill-a/stable-baselines","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ghostFaceKillah/expert","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/monte-ai/expert-augmented-acktr","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/openai/baselines","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/saschaschramm/Pong","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/jrobine/actor-critic","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ermongroup/higher_order_invariance","reach":{"status":"ok","spdx":"GPL-3.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/angelolovatto/raylab","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_draft_wrong":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"a44574392d168384","entry":"register","repo":"openai/baselines","repo_kind":"official","path":"baselines/common/models.py","file_url":"https://github.com/openai/baselines/blob/HEAD/baselines/common/models.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a44574392d168384"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}