{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/contrasting-exploration-in-parameter-and","title":"Contrasting Exploration in Parameter and Action Space: A Zeroth-Order Optimization Perspective","arxiv_id":"1901.11503","date":"2019-01-31","proceeding":null,"authors":["Anirudh Vemula","Wen Sun","J. Andrew Bagnell"],"abstract":"Black-box optimizers that explore in parameter space have often been shown to\noutperform more sophisticated action space exploration methods developed\nspecifically for the reinforcement learning problem. We examine these black-box\nmethods closely to identify situations in which they are worse than action\nspace exploration methods and those in which they are superior. Through simple\ntheoretical analyses, we prove that complexity of exploration in parameter\nspace depends on the dimensionality of parameter space, while complexity of\nexploration in action space depends on both the dimensionality of action space\nand horizon length. This is also demonstrated empirically by comparing simple\nexploration methods on several model problems, including Contextual Bandit,\nLinear Regression and Reinforcement Learning in continuous control.","url_abs":"http://arxiv.org/abs/1901.11503v1","url_pdf":"http://arxiv.org/pdf/1901.11503v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"contrasting-exploration-in-parameter-and","repo_url":"https://github.com/LAIRLAB/contrasting_exploration_rl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"},{"task_slug":"regression-1","task_name":"regression"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1901.11503","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}