{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-improving-deep-reinforcement-learning-for","title":"On Improving Deep Reinforcement Learning for POMDPs","arxiv_id":"1704.07978","date":"2017-04-26","proceeding":null,"authors":["Pengfei Zhu","Xin Li","Pascal Poupart","Guanghui Miao"],"abstract":"Deep Reinforcement Learning (RL) recently emerged as one of the most\ncompetitive approaches for learning in sequential decision making problems with\nfully observable environments, e.g., computer Go. However, very little work has\nbeen done in deep RL to handle partially observable environments. We propose a\nnew architecture called Action-specific Deep Recurrent Q-Network (ADRQN) to\nenhance learning performance in partially observable domains. Actions are\nencoded by a fully connected layer and coupled with a convolutional observation\nto form an action-observation pair. The time series of action-observation pairs\nare then integrated by an LSTM layer that learns latent states based on which a\nfully connected layer computes Q-values as in conventional Deep Q-Networks\n(DQNs). We demonstrate the effectiveness of our new architecture in several\npartially observable domains, including flickering Atari games.","url_abs":"http://arxiv.org/abs/1704.07978v6","url_pdf":"http://arxiv.org/pdf/1704.07978v6.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-improving-deep-reinforcement-learning-for","repo_url":"https://github.com/bit1029public/ADRQN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"atari-games","task_name":"Atari Games"},{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"sequential-decision-making","task_name":"Sequential Decision Making"},{"task_slug":"time-series-1","task_name":"Time Series"},{"task_slug":"time-series","task_name":"Time Series Analysis"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.07978","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}