Methods › Reinforcement Learning › Randomized Value Functions › REM

Random Ensemble Mixture

REM

48 papers tagged archive 2025-07-28

Introduced by Rishabh Agarwal et al. in An Optimistic Perspective on Offline Reinforcement Learning

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

Random Ensemble Mixture (REM) is an easy to implement extension of DQN inspired by Dropout. The key intuition behind REM is that if one has access to multiple estimates of Q-values, then a weighted combination of the Q-value estimates is also an estimate for Q-values. Accordingly, in each training step, REM randomly combines multiple Q-value estimates and uses this random combination for robust training.

PaperSource

Papers archive 2025-07-28

30 shown of 48, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 69 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
EEG6
Reinforcement Learning (RL)6
Reinforcement Learning5
reinforcement-learning5
DQN Replay Dataset3
Data Augmentation3
Offline RL3
Atari Games2
Classification2
Computational Efficiency2
Continual Learning2
Diversity2
Management2
Object2
Q-Learning2
Relation2
Sleep Staging2
Super-Resolution2
Ad-Hoc Information Retrieval1
Additive models1

Usage over time archive 2025-07-28

Papers per year tagged with REM: 2019 to 2025, peak 14 14 0 2019: 1 paper 2019 2020: 4 papers 2020 2021: 11 papers 2021 2022: 5 papers 2022 2023: 11 papers 2023 2024: 14 papers 2024 2025: 2 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (48 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Randomized Value FunctionsOff-Policy TD ControlQ-Learning Networks

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections