{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/wall-e-an-efficient-reinforcement-learning","title":"WALL-E: An Efficient Reinforcement Learning Research Framework","arxiv_id":"1901.06086","date":"2019-01-18","proceeding":null,"authors":["Tianbing Xu","Andrew Zhang","Liang Zhao"],"abstract":"There are two halves to RL systems: experience collection time and policy\nlearning time. For a large number of samples in rollouts, experience collection\ntime is the major bottleneck. Thus, it is necessary to speed up the rollout\ngeneration time with multi-process architecture support. Our work, dubbed\nWALL-E, utilizes multiple rollout samplers running in parallel to rapidly\ngenerate experience. Due to our parallel samplers, we experience not only\nfaster convergence times, but also higher average reward thresholds. For\nexample, on the MuJoCo HalfCheetah-v2 task, with $N = 10$ parallel sampler\nprocesses, we are able to achieve much higher average return than those from\nusing only a single process architecture.","url_abs":"http://arxiv.org/abs/1901.06086v2","url_pdf":"http://arxiv.org/pdf/1901.06086v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"wall-e-an-efficient-reinforcement-learning","repo_url":"https://github.com/harrybraviner/self_directed_rl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"mujoco","task_name":"MuJoCo"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}