Methods › Reinforcement Learning › Value Function Estimation › N-step Returns

N-step Returns

29 papers tagged archive 2025-07-28

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

n-step Returns are used for value function estimation in reinforcement learning. Specifically, for n steps we can write the complete return as:

Rₜ⁽ⁿ⁾ = rₜ₊₁ + γrₜ₊₂ + ⋯+ γⁿ⁻¹ₜ₊ₙ + γⁿVₜ(sₜ₊ₙ)

We can then write an n-step backup, in the style of TD learning, as:

ΔVᵣ(sₜ) = α[Rₜ⁽ⁿ⁾ - Vₜ(sₜ)]

Multi-step returns often lead to faster learning with suitably tuned n.

Image Credit: Sutton and Barto, Reinforcement Learning

Papers archive 2025-07-28

29 shown of 29, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 30 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Reinforcement Learning (RL)21
reinforcement-learning15
Deep Reinforcement Learning13
Reinforcement Learning12
Atari Games6
Q-Learning5
Continuous Control4
continuous-control4
Decision Making3
OpenAI Gym3
Benchmarking2
Distributional Reinforcement Learning2
BIG-bench Machine Learning1
Chunking1
Computational Efficiency1
DQN Replay Dataset1
Diversity1
Efficient Exploration1
Ensemble Learning1
Game of Go1

Usage over time archive 2025-07-28

Papers per year tagged with N-step Returns: 2017 to 2025, peak 7 7 0 2017: 2 papers 2017 2018: 3 papers 2018 2019: 2 papers 2019 2020: 7 papers 2020 2021: 4 papers 2021 2022: 4 papers 2022 2023: 3 papers 2023 2024: 2 papers 2024 2025: 2 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (29 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Value Function Estimation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections