{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ace-an-actor-ensemble-algorithm-for","title":"ACE: An Actor Ensemble Algorithm for Continuous Control with Tree Search","arxiv_id":"1811.02696","date":"2018-11-06","proceeding":null,"authors":["Shangtong Zhang","Hao Chen","Hengshuai Yao"],"abstract":"In this paper, we propose an actor ensemble algorithm, named ACE, for\ncontinuous control with a deterministic policy in reinforcement learning. In\nACE, we use actor ensemble (i.e., multiple actors) to search the global maxima\nof the critic. Besides the ensemble perspective, we also formulate ACE in the\noption framework by extending the option-critic architecture with deterministic\nintra-option policies, revealing a relationship between ensemble and options.\nFurthermore, we perform a look-ahead tree search with those actors and a\nlearned value prediction model, resulting in a refined value estimation. We\ndemonstrate a significant performance boost of ACE over DDPG and its variants\nin challenging physical robot simulators.","url_abs":"http://arxiv.org/abs/1811.02696v1","url_pdf":"http://arxiv.org/pdf/1811.02696v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ace-an-actor-ensemble-algorithm-for","repo_url":"https://github.com/ShangtongZhang/DeepRL","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"value-prediction","task_name":"Value prediction"},{"task_slug":"continuous-control","task_name":"continuous-control"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"adam","method_name":"Adam"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"ddpg","method_name":"DDPG"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"experience-replay","method_name":"Experience Replay"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"weight-decay","method_name":"Weight Decay"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1811.02696","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}