Methods › Reinforcement Learning

Reinforcement Learning

105 methods 28 collections 5,935 papers tagged archive 2025-07-28

Collections are ordered by tagged papers; each shows its 5 most-tagged methods.

Off-Policy TD Control

5 methods · 1,854 papers

Q-Learning

1,734 papers

Double Q-learning

112 papers

REM

Random Ensemble Mixture

48 papers

Expected Sarsa

9 papers

Policy Gradient Methods

23 methods · 1,633 papers

PPO

Proximal Policy Optimization

949 papers

DDPG

Deep Deterministic Policy Gradient

218 papers

REINFORCE

185 papers

TD3

Twin Delayed Deep Deterministic

117 papers

A2C

82 papers

5 shown of 23 methods →

Replay Memory

2 methods · 933 papers

Experience Replay

865 papers

Value Function Estimation

6 methods · 610 papers

HOC

High-Order Consensuses

515 papers

V-trace

34 papers

Retrace

31 papers

N-step Returns

29 papers

5 shown of 6 methods →

Q-Learning Networks

9 methods · 532 papers

DQN

Deep Q-Network

519 papers

REM

Random Ensemble Mixture

48 papers

Double DQN

45 papers

Dueling Network

23 papers

Rainbow DQN

9 papers

5 shown of 9 methods →

Offline Reinforcement Learning Methods

5 methods · 485 papers

DPO

Direct Preference Optimization

409 papers

URL

Umbrella Reinforcement Learning

34 papers

R2D2

Recurrent Replay Distributed DQN

25 papers

IQL

Implicit Q-Learning

16 papers

DeepCubeAI

DeepCubeA + Imagination

1 paper

Heuristic Search Algorithms

9 methods · 453 papers

GA

Genetic Algorithms

259 papers

Firefly algorithm

24 papers

4D A*

Four-dimensional A-star

1 paper

5 shown of 9 methods →

Video Game Models

2 methods · 432 papers

CARLA

CARLA: An Open Urban Driving Simulator

422 papers

AlphaStar

DeepMind AlphaStar

10 papers

Exploration Strategies

2 methods · 401 papers

Counterfactuals

Counterfactuals Explanations

400 papers

gSDE

Generalized State-Dependent Exploration

1 paper

Reinforcement Learning Frameworks

10 methods · 323 papers

AM

Attention Model

274 papers

RLAIF

Reinforcement Learning from AI Feedback

19 papers

SCST

Self-critical Sequence Training

12 papers

POMO

6 papers

TLA

Temporally Layered Architecture

6 papers

5 shown of 10 methods →

Board Game Models

3 methods · 155 papers

AlphaZero

114 papers

MuZero

46 papers

TD-Gammon

4 papers

On-Policy TD Control

5 methods · 70 papers

Sarsa

56 papers

TD Lambda

14 papers

Expected Sarsa

9 papers

Sarsa Lambda

0 papers

Randomized Value Functions

2 methods · 59 papers

REM

Random Ensemble Mixture

48 papers

Noisy Linear Layer

11 papers

Distributed Reinforcement Learning

6 methods · 39 papers

IMPALA

16 papers

Ape-X

10 papers

DD-PPO

Decentralized Distributed Proximal Policy Optimization

8 papers

APPO

Asynchronous Proximal Policy Optimization

4 papers

SEED RL

2 papers

5 shown of 6 methods →

Eligibility Traces

3 methods · 26 papers

Eligibility Trace

11 papers

Behaviour Policies

2 methods · 21 papers

Go-Explore

16 papers

Imitation Learning Methods

5 methods · 12 papers

IQ-Learn

Inverse Q-Learning

4 papers

CLIPort

3 papers

PWIL

Primal Wasserstein Imitation Learning

3 papers

CILO

Continuous Imitation Learning from Observation

1 paper

b2b transfer learning

building to building transfer learning

1 paper

Efficient Planning

1 method · 6 papers

State Similarity Metrics

2 methods · 5 papers

Card Game Models

1 method · 4 papers

DouZero

4 papers

RL Transformers

2 methods · 3 papers

GTrXL

Gated Transformer-XL

3 papers

CoBERL

Contrastive BERT

1 paper

Actor-Critic Algorithms

1 method · 1 paper

FORK

Forward-Looking Actor

1 paper

Bayesian Reinforcement Learning

1 method · 1 paper

Bayesian REX

Bayesian Reward Extrapolation

1 paper

Density Ratio Learning

1 method · 1 paper

GradientDICE

1 paper

Environment Design Methods

1 method · 1 paper

Motion Control

1 method · 1 paper

PPMC

Path Planning and Motion Control

1 paper

Path Planning

1 method · 1 paper

PPMC

Path Planning and Motion Control

1 paper

Policy Evaluation

1 method · 1 paper

KOVA

Kalman Optimization for Value Approximation

1 paper