{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bridging-the-gap-between-value-and-policy","title":"Bridging the Gap Between Value and Policy Based Reinforcement Learning","arxiv_id":"1702.08892","date":"2017-02-28","proceeding":"NeurIPS 2017 12","authors":["Ofir Nachum","Mohammad Norouzi","Kelvin Xu","Dale Schuurmans"],"abstract":"We establish a new connection between value and policy based reinforcement\nlearning (RL) based on a relationship between softmax temporal value\nconsistency and policy optimality under entropy regularization. Specifically,\nwe show that softmax consistent action values correspond to optimal entropy\nregularized policy probabilities along any action sequence, regardless of\nprovenance. From this observation, we develop a new RL algorithm, Path\nConsistency Learning (PCL), that minimizes a notion of soft consistency error\nalong multi-step action sequences extracted from both on- and off-policy\ntraces. We examine the behavior of PCL in different scenarios and show that PCL\ncan be interpreted as generalizing both actor-critic and Q-learning algorithms.\nWe subsequently deepen the relationship by showing how a single model can be\nused to represent both a policy and the corresponding softmax state values,\neliminating the need for a separate critic. The experimental evaluation\ndemonstrates that PCL significantly outperforms strong actor-critic and\nQ-learning baselines across several benchmarks.","url_abs":"http://arxiv.org/abs/1702.08892v3","url_pdf":"http://arxiv.org/pdf/1702.08892v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bridging-the-gap-between-value-and-policy","repo_url":"https://github.com/tensorflow/models","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"q-learning","task_name":"Q-Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[{"method_slug":"q-learning","method_name":"Q-Learning"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1702.08892","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}