Methods › Reinforcement Learning › Replay Memory › Experience Replay › Papers, page 9
Experience Replay
Papers archive 2025-07-28
archive papers tagged: 865 · with a code link: 317 · where Syntology ran a sample: 94 (86 with a run with no instrument failure, 8 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (94 of 865 tagged: 86 with a run with no instrument failure, 8 where every run was a failure of Syntology's instrument)
Page 9 of 9: papers 801 to 865 of 865, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Dynamic Weights in Multi-Objective Deep Reinforcement Learning 20 Sep 2018 · 3 repositories · arXiv:1809.07803
-
Curriculum goal masking for continuous deep reinforcement learning 17 Sep 2018 · 0 repositories · arXiv:1809.06146
-
Generalizing Across Multi-Objective Reward Functions in Deep Reinforcement Learning 17 Sep 2018 · 1 repository · arXiv:1809.06364Syntology 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
Improvements on Hindsight Learning 16 Sep 2018 · 0 repositories · arXiv:1809.06719
-
Learning Adaptive Display Exposure for Real-Time Advertising 10 Sep 2018 · 0 repositories · arXiv:1809.03149
-
ARCHER: Aggressive Rewards to Counter bias in Hindsight Experience Replay 6 Sep 2018 · 1 repository · arXiv:1809.02070
-
Adversarial Deep Reinforcement Learning in Portfolio Management 29 Aug 2018 · 5 repositories · arXiv:1808.09940
-
Goal-oriented Dialogue Policy Learning from Failures 20 Aug 2018 · 0 repositories · arXiv:1808.06497
-
Bipedal Walking Robot using Deep Deterministic Policy Gradient 16 Jul 2018 · 3 repositories · arXiv:1807.05924
-
Remember and Forget for Experience Replay 16 Jul 2018 · 2 repositories · arXiv:1807.05827Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Deterministic Policy Gradients With General State Transitions 10 Jul 2018 · 0 repositories · arXiv:1807.03708
-
Learning to Explore via Meta-Policy Gradient 1 Jul 2018 · 0 repositories
-
Stroke-based Character Reconstruction 23 Jun 2018 · 1 repository · arXiv:1806.08990
-
Organizing Experience: A Deeper Look at Replay Mechanisms for Sample-based Planning in Continuous State Domains 12 Jun 2018 · 0 repositories · arXiv:1806.04624
-
Randomized Value Functions via Multiplicative Normalizing Flows 6 Jun 2018 · 2 repositories · arXiv:1806.02315
-
Sample-Efficient Deep Reinforcement Learning via Episodic Backward Update 31 May 2018 · 1 repository · arXiv:1805.12375
-
Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models 30 May 2018 · 9 repositories · arXiv:1805.12114Syntology community repositories only · 12 ran (of which 5 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 5 where Syntology's instrument failed) · 7 unverified (of 19 harvested samples) · 17 pointer-only (licence)
-
Advances in Experience Replay 15 May 2018 · 1 repository · arXiv:1805.05536
-
Do deep reinforcement learning agents model intentions? 15 May 2018 · 1 repository · arXiv:1805.06020
-
Metatrace Actor-Critic: Online Step-size Tuning by Meta-gradient Descent for Reinforcement Learning Control 10 May 2018 · 0 repositories · arXiv:1805.04514
-
Multiagent Soft Q-Learning 25 Apr 2018 · 0 repositories · arXiv:1804.09817
-
State Distribution-aware Sampling for Deep Q-learning 23 Apr 2018 · 0 repositories · arXiv:1804.08619
-
Learning to Explore with Meta-Policy Gradient 13 Mar 2018 · 0 repositories · arXiv:1803.05044
-
Distributed Prioritized Experience Replay 2 Mar 2018 · 15 repositories · arXiv:1803.00933Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 15 harvested samples) · 3 pointer-only (licence)
-
Addressing Function Approximation Error in Actor-Critic Methods 26 Feb 2018 · 67 repositories · arXiv:1802.09477Syntology community repositories only · 26 ran (of which 0 constructed an object rather than computing a result; 25 with no instrument failure: 1 honoured, 1 violated, 23 with no contract checked; 1 where Syntology's instrument failed) · 10 unverified (of 36 harvested samples) · 21 pointer-only (licence)
-
Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research 26 Feb 2018 · 28 repositories · arXiv:1802.09464
-
Weighted Double Deep Multiagent Reinforcement Learning in Stochastic Cooperative Environments 23 Feb 2018 · 0 repositories · arXiv:1802.08534
-
Continual Reinforcement Learning with Complex Synapses 20 Feb 2018 · 0 repositories · arXiv:1802.07239
-
A Deep Q-Learning Agent for the L-Game with Variable Batch Training 17 Feb 2018 · 1 repository · arXiv:1802.06225
-
GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms 14 Feb 2018 · 1 repository · arXiv:1802.05054Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Efficient Exploration through Bayesian Deep Q-Networks 13 Feb 2018 · 1 repository · arXiv:1802.04412
-
Sample Efficient Deep Reinforcement Learning for Dialogue Systems with Large Action Spaces 11 Feb 2018 · 0 repositories · arXiv:1802.03753
-
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures 5 Feb 2018 · 24 repositories · arXiv:1802.01561Syntology community repositories only · 16 ran (of which 6 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 1 violated, 8 with no contract checked; 6 where Syntology's instrument failed) · 18 unverified (of 34 harvested samples) · 3 pointer-only (licence)
-
Pretraining Deep Actor-Critic Reinforcement Learning Algorithms With Expert Demonstrations 31 Jan 2018 · 0 repositories · arXiv:1801.10459
-
Deep In-GPU Experience Replay 9 Jan 2018 · 0 repositories · arXiv:1801.03138
-
Faster Deep Q-learning using Neural Episodic Control 6 Jan 2018 · 0 repositories · arXiv:1801.01968
-
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor 4 Jan 2018 · 86 repositories · arXiv:1801.01290Syntology community repositories only · 91 ran (of which 61 constructed an object rather than computing a result; 73 with no instrument failure: 3 honoured, 1 violated, 69 with no contract checked; 18 where Syntology's instrument failed) · 57 unverified (of 148 harvested samples) · 66 pointer-only (licence)
-
ScreenerNet: Learning Self-Paced Curriculum for Deep Neural Networks 3 Jan 2018 · 0 repositories · arXiv:1801.00904
-
ViZDoom: DRQN with Prioritized Experience Replay, Double-Q Learning, & Snapshot Ensembling 3 Jan 2018 · 0 repositories · arXiv:1801.01000
-
Learning to Run with Actor-Critic Ensemble 25 Dec 2017 · 2 repositories · arXiv:1712.08987
-
A Deeper Look at Experience Replay 4 Dec 2017 · 4 repositories · arXiv:1712.01275
-
AMBER: Adaptive Multi-Batch Experience Replay for Continuous Action Control 12 Oct 2017 · 0 repositories · arXiv:1710.04423
-
A novel DDPG method with prioritized experience replay 1 Oct 2017 · 1 repository
-
Overcoming Exploration in Reinforcement Learning with Demonstrations 28 Sep 2017 · 3 repositories · arXiv:1709.10089Syntology 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Linear Stochastic Approximation: Constant Step-Size and Iterate Averaging 12 Sep 2017 · 0 repositories · arXiv:1709.04073
-
Improving Stochastic Policy Gradients in Continuous Control with Deep Reinforcement Learning using the Beta Distribution 1 Aug 2017 · 0 repositories
-
Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards 27 Jul 2017 · 4 repositories · arXiv:1707.08817
-
Lenient Multi-Agent Deep Reinforcement Learning 14 Jul 2017 · 1 repository · arXiv:1707.04402
-
The Intentional Unintentional Agent: Learning to Solve Many Continuous Control Tasks Simultaneously 11 Jul 2017 · 0 repositories · arXiv:1707.03300
-
Hindsight Experience Replay 5 Jul 2017 · 28 repositories · arXiv:1707.01495Syntology 16 ran (of which 10 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 16 harvested samples) · 9 pointer-only (licence)
-
Sample-efficient Actor-Critic Reinforcement Learning with Supervised Data for Dialogue Management 1 Jul 2017 · 0 repositories · arXiv:1707.00130
-
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments 7 Jun 2017 · 86 repositories · arXiv:1706.02275Syntology official: no sample here; runs from other or unrecorded repositories · 75 ran (of which 54 constructed an object rather than computing a result; 68 with no instrument failure: 2 honoured, 0 violated, 66 with no contract checked; 7 where Syntology's instrument failed) · 68 unverified (of 143 harvested samples) · 99 pointer-only (licence)
-
Parameter Space Noise for Exploration 6 Jun 2017 · 10 repositories · arXiv:1706.01905Syntology 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Continual Learning with Deep Generative Replay 24 May 2017 · 5 repositories · arXiv:1705.08690
-
Discrete Sequential Prediction of Continuous Actions for Deep RL 14 May 2017 · 0 repositories · arXiv:1705.05035
-
Stabilising Experience Replay for Deep Multi-Agent Reinforcement Learning 28 Feb 2017 · 5 repositories · arXiv:1702.08887
-
Sample-efficient Deep Reinforcement Learning for Dialog Control 18 Dec 2016 · 0 repositories · arXiv:1612.06000
-
Sample Efficient Actor-Critic with Experience Replay 3 Nov 2016 · 7 repositories · arXiv:1611.01224Syntology 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Online Contrastive Divergence with Generative Replay: Experience Replay without Storing Data 18 Oct 2016 · 0 repositories · arXiv:1610.05555
-
Actor-critic versus direct policy search: a comparison based on sample complexity 29 Jun 2016 · 1 repository · arXiv:1606.09152
-
Continuous Deep Q-Learning with Model-based Acceleration 2 Mar 2016 · 8 repositories · arXiv:1603.00748Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Prioritized Experience Replay 18 Nov 2015 · 77 repositories · arXiv:1511.05952Syntology 80 ran (of which 62 constructed an object rather than computing a result; 72 with no instrument failure: 4 honoured, 0 violated, 68 with no contract checked; 8 where Syntology's instrument failed) · 31 unverified (of 111 harvested samples) · 43 pointer-only (licence)
-
Deep Reinforcement Learning with Double Q-learning 22 Sep 2015 · 97 repositories · arXiv:1509.06461Syntology 56 ran (of which 38 constructed an object rather than computing a result; 55 with no instrument failure: 0 honoured, 0 violated, 55 with no contract checked; 1 where Syntology's instrument failed) · 50 unverified (of 106 harvested samples) · 57 pointer-only (licence)
-
Continuous control with deep reinforcement learning 9 Sep 2015 · 161 repositories · arXiv:1509.02971Syntology 163 ran (of which 126 constructed an object rather than computing a result; 152 with no instrument failure: 3 honoured, 0 violated, 149 with no contract checked; 11 where Syntology's instrument failed) · 143 unverified (of 306 harvested samples) · 163 pointer-only (licence)
-
Playing Atari with Deep Reinforcement Learning 19 Dec 2013 · 112 repositories · arXiv:1312.5602Syntology 64 ran (of which 24 constructed an object rather than computing a result; 46 with no instrument failure: 5 honoured, 0 violated, 41 with no contract checked; 18 where Syntology's instrument failed) · 53 unverified (of 117 harvested samples) · 56 pointer-only (licence)