{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/where-did-my-optimum-go-an-empirical-analysis","title":"Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods","arxiv_id":"1810.02525","date":"2018-10-05","proceeding":null,"authors":["Peter Henderson","Joshua Romoff","Joelle Pineau"],"abstract":"Recent analyses of certain gradient descent optimization methods have shown\nthat performance can degrade in some settings - such as with stochasticity or\nimplicit momentum. In deep reinforcement learning (Deep RL), such optimization\nmethods are often used for training neural networks via the temporal difference\nerror or policy gradient. As an agent improves over time, the optimization\ntarget changes and thus the loss landscape (and local optima) change. Due to\nthe failure modes of those methods, the ideal choice of optimizer for Deep RL\nremains unclear. As such, we provide an empirical analysis of the effects that\na wide range of gradient descent optimizers and their hyperparameters have on\npolicy gradient methods, a subset of Deep RL algorithms, for benchmark\ncontinuous control tasks. We find that adaptive optimizers have a narrow window\nof effective learning rates, diverging in other cases, and that the\neffectiveness of momentum varies depending on the properties of the\nenvironment. Our analysis suggests that there is significant interplay between\nthe dynamics of the environment and Deep RL algorithm properties which aren't\nnecessarily accounted for by traditional adaptive gradient methods. We provide\nsuggestions for optimal settings of current methods and further lines of\nresearch based on our findings.","url_abs":"http://arxiv.org/abs/1810.02525v1","url_pdf":"http://arxiv.org/pdf/1810.02525v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"where-did-my-optimum-go-an-empirical-analysis","repo_url":"https://github.com/facebookresearch/WhereDidMyOptimumGo","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"policy-gradient-methods","task_name":"Policy Gradient Methods"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"continuous-control","task_name":"continuous-control"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.02525","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1810.02525"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/facebookresearch/WhereDidMyOptimumGo","reach":null}],"summary":{"ran_fixture":2,"ran_honours":1},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"e386eb38acef2cda","entry":"fix_point","repo":"facebookresearch/WhereDidMyOptimumGo","repo_kind":"official","path":"scripts/plot.py","file_url":"https://github.com/facebookresearch/WhereDidMyOptimumGo/blob/HEAD/scripts/plot.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"e386eb38acef2cda"}},{"code_sha256_prefix":"5dccd635d705075c","entry":"lim_by_game","repo":"facebookresearch/WhereDidMyOptimumGo","repo_kind":"official","path":"scripts/plot.py","file_url":"https://github.com/facebookresearch/WhereDidMyOptimumGo/blob/HEAD/scripts/plot.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"5dccd635d705075c"}},{"code_sha256_prefix":"5a8136a358eab81a","entry":"smooth_reward_curve","repo":"facebookresearch/WhereDidMyOptimumGo","repo_kind":"official","path":"scripts/plot.py","file_url":"https://github.com/facebookresearch/WhereDidMyOptimumGo/blob/HEAD/scripts/plot.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"5a8136a358eab81a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}