{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/memory-based-control-with-recurrent-neural","title":"Memory-based control with recurrent neural networks","arxiv_id":"1512.04455","date":"2015-12-14","proceeding":null,"authors":["Nicolas Heess","Jonathan J. Hunt","Timothy P. Lillicrap","David Silver"],"abstract":"Partially observed control problems are a challenging aspect of reinforcement\nlearning. We extend two related, model-free algorithms for continuous control\n-- deterministic policy gradient and stochastic value gradient -- to solve\npartially observed domains using recurrent neural networks trained with\nbackpropagation through time.\n  We demonstrate that this approach, coupled with long-short term memory is\nable to solve a variety of physical control problems exhibiting an assortment\nof memory requirements. These include the short-term integration of information\nfrom noisy sensors and the identification of system parameters, as well as\nlong-term memory problems that require preserving information over many time\nsteps. We also demonstrate success on a combined exploration and memory problem\nin the form of a simplified version of the well-known Morris water maze task.\nFinally, we show that our approach can deal with high-dimensional observations\nby learning directly from pixels.\n  We find that recurrent deterministic and stochastic policies are able to\nlearn similarly good solutions to these tasks, including the water maze where\nthe agent must learn effective search strategies.","url_abs":"http://arxiv.org/abs/1512.04455v1","url_pdf":"http://arxiv.org/pdf/1512.04455v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"memory-based-control-with-recurrent-neural","repo_url":"https://github.com/fshamshirdar/pytorch-rdpg","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"memory-based-control-with-recurrent-neural","repo_url":"https://github.com/quantumiracle/Popular-RL-Algorithms","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"memory-based-control-with-recurrent-neural","repo_url":"https://github.com/stevenpjg/RDPG","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"continuous-control","task_name":"continuous-control"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1512.04455","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}