{"url":"/method/retrace","slug":"retrace","name":"Retrace","full_name":"Retrace","full_name_withheld":false,"description_markdown":"**Retrace** is an off-policy Q-value estimation algorithm which has guaranteed convergence for a target and behaviour policy $\\left(\\pi, \\beta\\right)$. With off-policy rollout for TD learning, we must use importance sampling for the update:\r\n\r\n$$ \\Delta{Q}^{\\text{imp}}\\left(S\\_{t}, A\\_{t}\\right) = \\gamma^{t}\\prod\\_{1\\leq{\\tau}\\leq{t}}\\frac{\\pi\\left(A\\_{\\tau}\\mid{S\\_{\\tau}}\\right)}{\\beta\\left(A\\_{\\tau}\\mid{S\\_{\\tau}}\\right)}\\delta\\_{t} $$\r\n\r\nThis product term can lead to high variance, so Retrace modifies $\\Delta{Q}$ to have importance weights truncated by no more than a constant $c$:\r\n\r\n$$ \\Delta{Q}^{\\text{imp}}\\left(S\\_{t}, A\\_{t}\\right) = \\gamma^{t}\\prod\\_{1\\leq{\\tau}\\leq{t}}\\min\\left(c, \\frac{\\pi\\left(A\\_{\\tau}\\mid{S\\_{\\tau}}\\right)}{\\beta\\left(A\\_{\\tau}\\mid{S\\_{\\tau}}\\right)}\\right)\\delta\\_{t} $$","description_state":"present","introduced_year":null,"introduced_by":{"title":"Safe and Efficient Off-Policy Reinforcement Learning","paper":"/paper/safe-and-efficient-off-policy-reinforcement","first_author":"Rémi Munos","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/safe-and-efficient-off-policy-reinforcement"},"source":{"url":"http://arxiv.org/abs/1606.02647v2","title":"Safe and Efficient Off-Policy Reinforcement Learning","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Reinforcement Learning","area_id":"reinforcement-learning","collection":"Value Function Estimation","url":"/methods/category/value-function-estimation","pwc_aliases":[]}],"n_papers_tagged":31,"archive_num_papers":31,"papers_newest_first":[{"paper":null,"title":"UI-Evol: Automatic Knowledge Evolving for Computer Use Agents","date":"2025-05-28","arxiv_id":"2505.21964","n_code_links":0,"syntology":null},{"paper":null,"title":"Generative Artificial Intelligence: Evolving Technology, Growing Societal Impact, and Opportunities for Information Systems Research","date":"2025-02-25","arxiv_id":"2503.05770","n_code_links":0,"syntology":null},{"paper":null,"title":"Dynamics of Resource Allocation in O-RANs: An In-depth Exploration of On-Policy and Off-Policy Deep Reinforcement Learning for Real-Time Applications","date":"2024-11-17","arxiv_id":"2412.01839","n_code_links":0,"syntology":null},{"paper":null,"title":"IDRetracor: Towards Visual Forensics Against Malicious Face Swapping","date":"2024-08-13","arxiv_id":"2408.06635","n_code_links":0,"syntology":null},{"paper":"/paper/joint-physical-digital-facial-attack","title":"Joint Physical-Digital Facial Attack Detection Via Simulating Spoofing Clues","date":"2024-04-12","arxiv_id":"2404.08450","n_code_links":3,"syntology":null},{"paper":null,"title":"Off-policy Distributional Q($λ$): Distributional RL without Importance Sampling","date":"2024-02-08","arxiv_id":"2402.05766","n_code_links":0,"syntology":null},{"paper":null,"title":"Network-thinking to optimize surveillance and control of crop parasites. A review","date":"2023-10-11","arxiv_id":"2310.07442","n_code_links":0,"syntology":null},{"paper":null,"title":"Distributional Estimation of Data Uncertainty for Surveillance Face Anti-spoofing","date":"2023-09-18","arxiv_id":"2309.09485","n_code_links":0,"syntology":null},{"paper":null,"title":"PDVN: A Patch-based Dual-view Network for Face Liveness Detection using Light Field Focal Stack","date":"2023-01-17","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"AcceRL: Policy Acceleration Framework for Deep Reinforcement Learning","date":"2022-11-28","arxiv_id":"2211.15023","n_code_links":0,"syntology":null},{"paper":null,"title":"Asynchronous Curriculum Experience Replay: A Deep Reinforcement Learning Approach for UAV Autonomous Motion Control in Unknown Dynamic Environments","date":"2022-07-04","arxiv_id":"2207.01251","n_code_links":0,"syntology":null},{"paper":null,"title":"Safe-FinRL: A Low Bias and Variance Deep Reinforcement Learning Implementation for High-Freq Stock Trading","date":"2022-06-13","arxiv_id":"2206.05910","n_code_links":0,"syntology":null},{"paper":null,"title":"Bias-inducing geometries: an exactly solvable data model with fairness implications","date":"2022-05-31","arxiv_id":"2205.15935","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep Learning with Logical Constraints","date":"2022-05-01","arxiv_id":"2205.00523","n_code_links":0,"syntology":null},{"paper":null,"title":"Is Word Error Rate a good evaluation metric for Speech Recognition in Indic Languages?","date":"2022-03-30","arxiv_id":"2203.16601","n_code_links":0,"syntology":null},{"paper":null,"title":"Marginalized Operators for Off-policy Reinforcement Learning","date":"2022-03-30","arxiv_id":"2203.16177","n_code_links":0,"syntology":null},{"paper":null,"title":"Improving the Efficiency of Off-Policy Reinforcement Learning by Accounting for Past Decisions","date":"2021-12-23","arxiv_id":"2112.12281","n_code_links":0,"syntology":null},{"paper":null,"title":"Learning Reward Machines: A Study in Partially Observable Reinforcement Learning","date":"2021-12-17","arxiv_id":"2112.09477","n_code_links":0,"syntology":null},{"paper":null,"title":"Human Languages with Greater Information Density Increase Communication Speed, but Decrease Conversation Breadth","date":"2021-12-15","arxiv_id":"2112.08491","n_code_links":0,"syntology":null},{"paper":null,"title":"Dynamics of the market states in the space of correlation matrices with applications to financial markets","date":"2021-07-12","arxiv_id":"2107.05663","n_code_links":0,"syntology":null},{"paper":"/paper/a-deeppixbis-attentional-angular-margin-for","title":"A-DeepPixBis: Attentional Angular Margin for Face Anti-Spoofing","date":"2021-03-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/re-satellite-image-time-series-classification","title":"[Re] Satellite Image Time Series Classification with Pixel-Set Encoders and Temporal Self-Attention","date":"2020-12-06","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Exploiting the potential of deep reinforcement learning for classification tasks in high-dimensional and unstructured data","date":"2019-12-20","arxiv_id":"1912.09595","n_code_links":0,"syntology":null},{"paper":"/paper/learning-reward-machines-for-partially","title":"Learning Reward Machines for Partially Observable Reinforcement Learning","date":"2019-12-01","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Gap-Increasing Policy Evaluation for Efficient and Noise-Tolerant Reinforcement Learning","date":"2019-06-18","arxiv_id":"1906.07586","n_code_links":0,"syntology":null},{"paper":"/paper/understanding-multi-step-deep-reinforcement","title":"Understanding Multi-Step Deep Reinforcement Learning: A Systematic Study of the DQN Target","date":"2019-01-22","arxiv_id":"1901.07510","n_code_links":1,"syntology":null},{"paper":null,"title":"Sample Efficient Deep Reinforcement Learning for Dialogue Systems with Large Action Spaces","date":"2018-02-11","arxiv_id":"1802.03753","n_code_links":0,"syntology":null},{"paper":null,"title":"Pretraining Deep Actor-Critic Reinforcement Learning Algorithms With Expert Demonstrations","date":"2018-01-31","arxiv_id":"1801.10459","n_code_links":0,"syntology":null},{"paper":"/paper/the-reactor-a-fast-and-sample-efficient-actor","title":"The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning","date":"2017-04-15","arxiv_id":"1704.04651","n_code_links":0,"syntology":null},{"paper":"/paper/sample-efficient-actor-critic-with-experience","title":"Sample Efficient Actor-Critic with Experience Replay","date":"2016-11-03","arxiv_id":"1611.01224","n_code_links":7,"syntology":{"ran":0,"of":1,"unverified":1,"pointer_only":1}}],"papers_shown":30,"tasks":[{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":13},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":12},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":12},{"task":"/task/deep-reinforcement-learning","name":"Deep Reinforcement Learning","papers":9},{"task":"/task/face-anti-spoofing","name":"Face Anti-Spoofing","papers":3},{"task":"/task/face-recognition","name":"Face Recognition","papers":3},{"task":"/task/atari-games","name":"Atari Games","papers":2},{"task":"/task/classification","name":"General Classification","papers":2},{"task":"/task/partially-observable-reinforcement-learning","name":"Partially Observable Reinforcement Learning","papers":2},{"task":"/task/problem-decomposition","name":"Problem Decomposition","papers":2},{"task":"/task/time-series-1","name":"Time Series","papers":2},{"task":"/task/time-series","name":"Time Series Analysis","papers":2},{"task":"/task/automatic-speech-recognition-2","name":"Automatic Speech Recognition","papers":1},{"task":"/task/automatic-speech-recognition","name":"Automatic Speech Recognition (ASR)","papers":1},{"task":"/task/benchmarking","name":"Benchmarking","papers":1},{"task":"/task/binary-classification","name":"Binary Classification","papers":1},{"task":"/task/continuous-control","name":"Continuous Control","papers":1},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":1},{"task":"/task/decision-making","name":"Decision Making","papers":1},{"task":"/task/deep-learning","name":"Deep Learning","papers":1}],"tasks_shown":20,"n_tasks":38,"usage_by_year":[{"year":"2016","papers":2},{"year":"2017","papers":1},{"year":"2018","papers":2},{"year":"2019","papers":4},{"year":"2020","papers":1},{"year":"2021","papers":5},{"year":"2022","papers":7},{"year":"2023","papers":3},{"year":"2024","papers":4},{"year":"2025","papers":2}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/retrace"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}