{"url":"/method/n-step-returns","slug":"n-step-returns","name":"N-step Returns","full_name":"N-step Returns","full_name_withheld":false,"description_markdown":"**$n$-step Returns** are used for value function estimation in reinforcement learning. Specifically, for $n$ steps we can write the complete return as:\r\n\r\n$$ R\\_{t}^{(n)} = r\\_{t+1} + \\gamma{r}\\_{t+2} + \\cdots + \\gamma^{n-1}\\_{t+n} + \\gamma^{n}V\\_{t}\\left(s\\_{t+n}\\right) $$\r\n\r\nWe can then write an $n$-step backup, in the style of TD learning, as:\r\n\r\n$$ \\Delta{V}\\_{r}\\left(s\\_{t}\\right) = \\alpha\\left[R\\_{t}^{(n)} - V\\_{t}\\left(s\\_{t}\\right)\\right] $$\r\n\r\nMulti-step returns often lead to faster learning with suitably tuned $n$.\r\n\r\nImage Credit: Sutton and Barto, Reinforcement Learning","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":null,"title":null,"url_on_a_paper_host":false},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Reinforcement Learning","area_id":"reinforcement-learning","collection":"Value Function Estimation","url":"/methods/category/value-function-estimation","pwc_aliases":[]}],"n_papers_tagged":29,"archive_num_papers":29,"papers_newest_first":[{"paper":"/paper/shapley-machine-a-game-theoretic-framework","title":"Shapley Machine: A Game-Theoretic Framework for N-Agent Ad Hoc Teamwork","date":"2025-06-12","arxiv_id":"2506.11285","n_code_links":1,"syntology":{"ran":3,"of":3,"unverified":0,"pointer_only":0}},{"paper":null,"title":"Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step Returns","date":"2025-03-05","arxiv_id":"2503.03660","n_code_links":0,"syntology":null},{"paper":"/paper/beyond-the-rainbow-high-performance-deep","title":"Beyond The Rainbow: High Performance Deep Reinforcement Learning on a Desktop PC","date":"2024-11-06","arxiv_id":"2411.03820","n_code_links":3,"syntology":{"ran":13,"of":25,"unverified":12,"pointer_only":21}},{"paper":null,"title":"Learning in complex action spaces without policy gradients","date":"2024-10-08","arxiv_id":"2410.06317","n_code_links":0,"syntology":null},{"paper":null,"title":"Mitigating Estimation Errors by Twin TD-Regularized Actor and Critic for Deep Reinforcement Learning","date":"2023-11-07","arxiv_id":"2311.03711","n_code_links":0,"syntology":null},{"paper":"/paper/sdgym-low-code-reinforcement-learning","title":"SDGym: Low-Code Reinforcement Learning Environments using System Dynamics Models","date":"2023-10-19","arxiv_id":"2310.12494","n_code_links":1,"syntology":null},{"paper":null,"title":"A Long $N$-step Surrogate Stage Reward for Deep Reinforcement Learning","date":"2023-09-21","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/reducing-variance-in-temporal-difference","title":"Reducing Variance in Temporal-Difference Value Estimation via Ensemble of Deep Networks","date":"2022-09-16","arxiv_id":"2209.07670","n_code_links":1,"syntology":null},{"paper":"/paper/dna-proximal-policy-optimization-with-a-dual","title":"DNA: Proximal Policy Optimization with a Dual Network Architecture","date":"2022-06-20","arxiv_id":"2206.10027","n_code_links":1,"syntology":null},{"paper":"/paper/gamma-and-vega-hedging-using-deep","title":"Gamma and Vega Hedging Using Deep Distributional Reinforcement Learning","date":"2022-05-10","arxiv_id":"2205.05614","n_code_links":1,"syntology":null},{"paper":"/paper/revisiting-gaussian-mixture-critic-in-off","title":"Revisiting Gaussian mixture critics in off-policy reinforcement learning: a sample-based approach","date":"2022-04-21","arxiv_id":"2204.10256","n_code_links":1,"syntology":null},{"paper":"/paper/deep-reinforcement-learning-at-the-edge-of","title":"Deep Reinforcement Learning at the Edge of the Statistical Precipice","date":"2021-08-30","arxiv_id":"2108.13264","n_code_links":3,"syntology":{"ran":5,"of":5,"unverified":0,"pointer_only":0}},{"paper":"/paper/a-coevolutionairy-approach-to-deep-multi","title":"A coevolutionary approach to deep multi-agent reinforcement learning","date":"2021-04-12","arxiv_id":"2104.05610","n_code_links":1,"syntology":null},{"paper":null,"title":"Adaptive N-step Bootstrapping with Off-policy Data","date":"2021-01-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Weighted Bellman Backups for Improved Signal-to-Noise in Q-Updates","date":"2021-01-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/tonic-a-deep-reinforcement-learning-library","title":"Tonic: A Deep Reinforcement Learning Library for Fast Prototyping and Benchmarking","date":"2020-11-15","arxiv_id":"2011.07537","n_code_links":1,"syntology":{"ran":3,"of":4,"unverified":1,"pointer_only":0}},{"paper":null,"title":"A New Approach for Tactical Decision Making in Lane Changing: Sample Efficient Deep Q Learning with a Safety Feedback Reward","date":"2020-09-24","arxiv_id":"2009.11905","n_code_links":0,"syntology":null},{"paper":"/paper/munchausen-reinforcement-learning","title":"Munchausen Reinforcement Learning","date":"2020-07-28","arxiv_id":"2007.14430","n_code_links":6,"syntology":{"ran":12,"of":22,"unverified":10,"pointer_only":0}},{"paper":"/paper/revisiting-fundamentals-of-experience-replay","title":"Revisiting Fundamentals of Experience Replay","date":"2020-07-13","arxiv_id":"2007.06700","n_code_links":2,"syntology":{"ran":0,"of":1,"unverified":1,"pointer_only":1}},{"paper":"/paper/sunrise-a-simple-unified-framework-for","title":"SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement Learning","date":"2020-07-09","arxiv_id":"2007.04938","n_code_links":1,"syntology":null},{"paper":null,"title":"Distributed Uplink Beamforming in Cell-Free Networks Using Deep Reinforcement Learning","date":"2020-06-26","arxiv_id":"2006.15138","n_code_links":0,"syntology":null},{"paper":null,"title":"Sample-based Distributional Policy Gradient","date":"2020-01-08","arxiv_id":"2001.02652","n_code_links":0,"syntology":null},{"paper":null,"title":"Generative Adversarial Imagination for Sample Efficient Deep Reinforcement Learning","date":"2019-04-30","arxiv_id":"1904.13255","n_code_links":0,"syntology":null},{"paper":"/paper/tf-replicator-distributed-machine-learning","title":"TF-Replicator: Distributed Machine Learning for Researchers","date":"2019-02-01","arxiv_id":"1902.00465","n_code_links":1,"syntology":{"ran":2,"of":3,"unverified":1,"pointer_only":0}},{"paper":"/paper/macro-action-selection-with-deep","title":"Macro action selection with deep reinforcement learning in StarCraft","date":"2018-12-02","arxiv_id":"1812.00336","n_code_links":1,"syntology":null},{"paper":"/paper/distributed-distributional-deterministic","title":"Distributed Distributional Deterministic Policy Gradients","date":"2018-04-23","arxiv_id":"1804.08617","n_code_links":5,"syntology":null},{"paper":"/paper/distributed-prioritized-experience-replay","title":"Distributed Prioritized Experience Replay","date":"2018-03-02","arxiv_id":"1803.00933","n_code_links":15,"syntology":{"ran":0,"of":15,"unverified":15,"pointer_only":0}},{"paper":"/paper/rainbow-combining-improvements-in-deep","title":"Rainbow: Combining Improvements in Deep Reinforcement Learning","date":"2017-10-06","arxiv_id":"1710.02298","n_code_links":34,"syntology":{"ran":2,"of":6,"unverified":4,"pointer_only":1}},{"paper":null,"title":"Learning to Mix n-Step Returns: Generalizing lambda-Returns for Deep Reinforcement Learning","date":"2017-05-21","arxiv_id":"1705.07445","n_code_links":0,"syntology":null}],"papers_shown":29,"tasks":[{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":21},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":15},{"task":"/task/deep-reinforcement-learning","name":"Deep Reinforcement Learning","papers":13},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":12},{"task":"/task/atari-games","name":"Atari Games","papers":6},{"task":"/task/q-learning","name":"Q-Learning","papers":5},{"task":"/task/continuous-control","name":"Continuous Control","papers":4},{"task":"/task/continuous-control","name":"continuous-control","papers":4},{"task":"/task/decision-making","name":"Decision Making","papers":3},{"task":"/task/openai-gym","name":"OpenAI Gym","papers":3},{"task":"/task/benchmarking","name":"Benchmarking","papers":2},{"task":"/task/distributional-reinforcement-learning","name":"Distributional Reinforcement Learning","papers":2},{"task":"/task/machine-learning","name":"BIG-bench Machine Learning","papers":1},{"task":"/task/chunking","name":"Chunking","papers":1},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":1},{"task":"/task/dqn-replay-dataset","name":"DQN Replay Dataset","papers":1},{"task":"/task/diversity","name":"Diversity","papers":1},{"task":"/task/efficient-exploration","name":"Efficient Exploration","papers":1},{"task":"/task/ensemble-learning","name":"Ensemble Learning","papers":1},{"task":"/task/game-of-go","name":"Game of Go","papers":1}],"tasks_shown":20,"n_tasks":30,"usage_by_year":[{"year":"2017","papers":2},{"year":"2018","papers":3},{"year":"2019","papers":2},{"year":"2020","papers":7},{"year":"2021","papers":4},{"year":"2022","papers":4},{"year":"2023","papers":3},{"year":"2024","papers":2},{"year":"2025","papers":2}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/n-step-returns"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}