{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-run-with-actor-critic-ensemble","title":"Learning to Run with Actor-Critic Ensemble","arxiv_id":"1712.08987","date":"2017-12-25","proceeding":null,"authors":["Zhewei Huang","Shuchang Zhou","BoEr Zhuang","Xinyu Zhou"],"abstract":"We introduce an Actor-Critic Ensemble(ACE) method for improving the\nperformance of Deep Deterministic Policy Gradient(DDPG) algorithm. At inference\ntime, our method uses a critic ensemble to select the best action from\nproposals of multiple actors running in parallel. By having a larger candidate\nset, our method can avoid actions that have fatal consequences, while staying\ndeterministic. Using ACE, we have won the 2nd place in NIPS'17 Learning to Run\ncompetition, under the name of \"Megvii-hzwer\".","url_abs":"http://arxiv.org/abs/1712.08987v1","url_pdf":"http://arxiv.org/pdf/1712.08987v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-run-with-actor-critic-ensemble","repo_url":"https://github.com/hzwer/NIPS2017-LearningToRun","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"learning-to-run-with-actor-critic-ensemble","repo_url":"https://github.com/megvii-research/NIPS2017-LearningToRunACE","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[],"methods":[{"method_slug":"adam","method_name":"Adam"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"ddpg","method_name":"DDPG"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"experience-replay","method_name":"Experience Replay"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"weight-decay","method_name":"Weight Decay"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1712.08987","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}