Methods › General › Regularization › Entropy Regularization › Papers, page 12
Entropy Regularization
Papers archive 2025-07-28
archive papers tagged: 1,128 · with a code link: 451 · where Syntology ran a sample: 156 (129 with a run with no instrument failure, 27 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (156 of 1,128 tagged: 129 with a run with no instrument failure, 27 where every run was a failure of Syntology's instrument)
Page 12 of 12: papers 1,101 to 1,128 of 1,128, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Brandom-ian view of Reinforcement Learning towards strong-AI 7 Mar 2018 · 0 repositories · arXiv:1803.02912
-
Variational Inference for Policy Gradient 21 Feb 2018 · 0 repositories · arXiv:1802.07833
-
Sample Efficient Deep Reinforcement Learning for Dialogue Systems with Large Action Spaces 11 Feb 2018 · 0 repositories · arXiv:1802.03753
-
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures 5 Feb 2018 · 24 repositories · arXiv:1802.01561Syntology community repositories only · 16 ran (of which 6 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 1 violated, 8 with no contract checked; 6 where Syntology's instrument failed) · 18 unverified (of 34 harvested samples) · 3 pointer-only (licence)
-
Pretraining Deep Actor-Critic Reinforcement Learning Algorithms With Expert Demonstrations 31 Jan 2018 · 0 repositories · arXiv:1801.10459
-
An Empirical Analysis of Proximal Policy Optimization with Kronecker-factored Natural Gradients 17 Jan 2018 · 0 repositories · arXiv:1801.05566
-
Exploring Deep Recurrent Models with Reinforcement Learning for Molecule Design 1 Jan 2018 · 0 repositories
-
Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement Learning 18 Dec 2017 · 12 repositories · arXiv:1712.06567Syntology 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
Natural Value Approximators: Learning when to Trust Past Estimates 1 Dec 2017 · 0 repositories
-
Teaching a Machine to Read Maps with Deep Reinforcement Learning 20 Nov 2017 · 1 repository · arXiv:1711.07479
-
CARLA: An Open Urban Driving Simulator 10 Nov 2017 · 1 repository · arXiv:1711.03938
-
AMBER: Adaptive Multi-Batch Experience Replay for Continuous Action Control 12 Oct 2017 · 0 repositories · arXiv:1710.04423
-
Sparse Markov Decision Processes with Causal Sparse Tsallis Entropy Regularization for Reinforcement Learning 19 Sep 2017 · 0 repositories · arXiv:1709.06293
-
Improving Search through A3C Reinforcement Learning based Conversational Agent 17 Sep 2017 · 0 repositories · arXiv:1709.05638
-
Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation 17 Aug 2017 · 8 repositories · arXiv:1708.05144Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
DARLA: Improving Zero-Shot Transfer in Reinforcement Learning 26 Jul 2017 · 1 repository · arXiv:1707.08475
-
Learning Transferable Architectures for Scalable Image Recognition 21 Jul 2017 · 17 repositories · arXiv:1707.07012Syntology 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Proximal Policy Optimization Algorithms 20 Jul 2017 · 188 repositories · arXiv:1707.06347Syntology 99 ran (of which 53 constructed an object rather than computing a result; 71 with no instrument failure: 7 honoured, 2 violated, 62 with no contract checked; 28 where Syntology's instrument failed) · 77 unverified (of 176 harvested samples) · 94 pointer-only (licence)
-
Noisy Networks for Exploration 30 Jun 2017 · 15 repositories · arXiv:1706.10295Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Learning to Factor Policies and Action-Value Functions: Factored Action Space Representations for Deep Reinforcement learning 20 May 2017 · 0 repositories · arXiv:1705.07269
-
Feature Control as Intrinsic Motivation for Hierarchical Reinforcement Learning 18 May 2017 · 1 repository · arXiv:1705.06769
-
Equivalence Between Policy Gradients and Soft Q-Learning 21 Apr 2017 · 0 repositories · arXiv:1704.06440
-
The Reactor: A fast and sample-efficient Actor-Critic agent for Reinforcement Learning 15 Apr 2017 · 0 repositories · arXiv:1704.04651
-
Tactics of Adversarial Attack on Deep Reinforcement Learning Agents 8 Mar 2017 · 0 repositories · arXiv:1703.06748
-
Improving Policy Gradient by Exploring Under-appreciated Rewards 28 Nov 2016 · 0 repositories · arXiv:1611.09321
-
Reinforcement Learning through Asynchronous Advantage Actor-Critic on a GPU 18 Nov 2016 · 3 repositories · arXiv:1611.06256
-
Sample Efficient Actor-Critic with Experience Replay 3 Nov 2016 · 7 repositories · arXiv:1611.01224Syntology 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Asynchronous Methods for Deep Reinforcement Learning 4 Feb 2016 · 70 repositories · arXiv:1602.01783Syntology 60 ran (of which 20 constructed an object rather than computing a result; 51 with no instrument failure: 2 honoured, 1 violated, 48 with no contract checked; 9 where Syntology's instrument failed) · 35 unverified (of 95 harvested samples) · 20 pointer-only (licence)