Methods › Reinforcement Learning › Policy Gradient Methods › REINFORCE › Papers, page 2
REINFORCE
Papers archive 2025-07-28
archive papers tagged: 185 · with a code link: 80 · where Syntology ran a sample: 24 (20 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (24 of 185 tagged: 20 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument)
Page 2 of 2: papers 101 to 185 of 185, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Robust Dialogue Utterance Rewriting as Sequence Tagging 29 Dec 2020 · 1 repository · arXiv:2012.14535
-
On (Emergent) Systematic Generalisation and Compositionality in Visual Referential Games with Straight-Through Gumbel-Softmax Estimator 19 Dec 2020 · 1 repository · arXiv:2012.10776
-
Inferring learning rules from animal decision-making 1 Dec 2020 · 1 repository
-
TaylorGAN: Neighbor-Augmented Policy Update Towards Sample-Efficient Natural Language Generation 1 Dec 2020 · 1 repository
-
TaylorGAN: Neighbor-Augmented Policy Update for Sample-Efficient Natural Language Generation 27 Nov 2020 · 1 repository · arXiv:2011.13527
-
Hindsight Network Credit Assignment 24 Nov 2020 · 0 repositories · arXiv:2011.12351
-
Efficient Neural Architecture Search for End-to-end Speech Recognition via Straight-Through Gradients 11 Nov 2020 · 1 repository · arXiv:2011.05649
-
Dirichlet policies for reinforced factor portfolios 10 Nov 2020 · 0 repositories · arXiv:2011.05381
-
Guided Dialogue Policy Learning without Adversarial Learning in the Loop 1 Nov 2020 · 1 repository
-
POMO: Policy Optimization with Multiple Optima for Reinforcement Learning 30 Oct 2020 · 3 repositories · arXiv:2010.16011Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Sample Efficient Reinforcement Learning with REINFORCE 22 Oct 2020 · 0 repositories · arXiv:2010.11364
-
Learning by Competition of Self-Interested Reinforcement Learning Agents 19 Oct 2020 · 1 repository · arXiv:2010.09770
-
How Does Supernet Help in Neural Architecture Search? 16 Oct 2020 · 0 repositories · arXiv:2010.08219
-
MAP Propagation Algorithm: Faster Learning with a Team of Reinforcement Learning Agents 15 Oct 2020 · 1 repository · arXiv:2010.07893
-
Variance-Reduced Off-Policy Memory-Efficient Policy Search 14 Sep 2020 · 0 repositories · arXiv:2009.06548
-
Improving Language Generation with Sentence Coherence Objective 7 Sep 2020 · 1 repository · arXiv:2009.06358
-
ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement Learning 4 Sep 2020 · 1 repository · arXiv:2009.02010Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Stochastic Fine-grained Labeling of Multi-state Sign Glosses for Continuous Sign Language Recognition 1 Aug 2020 · 1 repository
-
An operator view of policy gradient methods 19 Jun 2020 · 0 repositories · arXiv:2006.11266
-
Model-based Adversarial Meta-Reinforcement Learning 16 Jun 2020 · 1 repository · arXiv:2006.08875
-
AutoHAS: Efficient Hyperparameter and Architecture Search 5 Jun 2020 · 0 repositories · arXiv:2006.03656
-
The Importance of Prior Knowledge in Precise Multimodal Prediction 4 Jun 2020 · 0 repositories · arXiv:2006.02636
-
Jointly Learning Environments and Control Policies with Projected Stochastic Gradient Ascent 2 Jun 2020 · 1 repository · arXiv:2006.01738
-
Angle-based Search Space Shrinking for Neural Architecture Search 28 Apr 2020 · 1 repository · arXiv:2004.13431Syntology 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Attention Routing: track-assignment detailed routing using attention-based reinforcement learning 20 Apr 2020 · 0 repositories · arXiv:2004.09473
-
Guided Dialog Policy Learning without Adversarial Learning in the Loop 7 Apr 2020 · 1 repository · arXiv:2004.03267Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
A Better Variant of Self-Critical Sequence Training 22 Mar 2020 · 1 repository · arXiv:2003.09971
-
A Hybrid Stochastic Policy Gradient Algorithm for Reinforcement Learning 1 Mar 2020 · 1 repository · arXiv:2003.00430Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Estimating Gradients for Discrete Random Variables by Sampling without Replacement 14 Feb 2020 · 1 repository · arXiv:2002.06043
-
Black-Box Optimization with Local Generative Surrogates 11 Feb 2020 · 1 repository · arXiv:2002.04632
-
Unsupervised Program Synthesis for Images By Sampling Without Replacement 27 Jan 2020 · 0 repositories · arXiv:2001.10119
-
Robust Federated Learning Through Representation Matching and Adaptive Hyper-parameters 30 Dec 2019 · 0 repositories · arXiv:1912.13075
-
UNAS: Differentiable Architecture Search Meets Reinforcement Learning 16 Dec 2019 · 1 repository · arXiv:1912.07651
-
Neural Predictor for Neural Architecture Search 2 Dec 2019 · 2 repositories · arXiv:1912.00848Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
Scene Graph based Image Retrieval -- A case study on the CLEVR Dataset 3 Nov 2019 · 0 repositories · arXiv:1911.00850
-
Hidden State Guidance: Improving Image Captioning using An Image Conditioned Autoencoder 31 Oct 2019 · 0 repositories · arXiv:1910.14208
-
All-Action Policy Gradient Methods: A Numerical Integration Approach 21 Oct 2019 · 0 repositories · arXiv:1910.09093
-
Analyzing the Variance of Policy Gradient Estimators for the Linear-Quadratic Regulator 2 Oct 2019 · 0 repositories · arXiv:1910.01249
-
Deep Reinforcement Learning with Modulated Hebbian plus Q Network Architecture 21 Sep 2019 · 1 repository · arXiv:1909.09902
-
ER-AE: Differentially Private Text Generation for Authorship Anonymization 20 Jul 2019 · 2 repositories · arXiv:1907.08736
-
Bridging by Word: Image Grounded Vocabulary Construction for Visual Captioning 1 Jul 2019 · 1 repository
-
Direct Policy Gradients: Direct Optimization of Policies in Discrete Action Spaces 14 Jun 2019 · 0 repositories · arXiv:1906.06062
-
Exploiting Uncertainty of Loss Landscape for Stochastic Optimization 30 May 2019 · 1 repository · arXiv:1905.13200Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Recurrent Existence Determination Through Policy Optimization 29 May 2019 · 0 repositories · arXiv:1905.13551
-
Interpretable Neural Predictions with Differentiable Binary Variables 20 May 2019 · 1 repository · arXiv:1905.08160Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
AM-LFS: AutoML for Loss Function Search 17 May 2019 · 1 repository · arXiv:1905.07375
-
ARSM: Augment-REINFORCE-Swap-Merge Estimator for Gradient Backpropagation Through Categorical Variables 4 May 2019 · 1 repository · arXiv:1905.01413Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Beyond Games: Bringing Exploration to Robots in Real-world 1 May 2019 · 0 repositories
-
Posterior-regularized REINFORCE for Instance Selection in Distant Supervision 17 Apr 2019 · 1 repository · arXiv:1904.08051
-
Generation of Synthetic Electronic Medical Record Text 6 Dec 2018 · 0 repositories · arXiv:1812.02793
-
Top-K Off-Policy Correction for a REINFORCE Recommender System 6 Dec 2018 · 1 repository · arXiv:1812.02353
-
ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware 2 Dec 2018 · 23 repositories · arXiv:1812.00332Syntology official (archive's flag): 2 ran · 17 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 2 honoured, 0 violated, 12 with no contract checked; 3 where Syntology's instrument failed) · 10 unverified (of 27 harvested samples) · 4 pointer-only (licence)
-
Learning to Exploit Stability for 3D Scene Parsing 1 Dec 2018 · 0 repositories
-
Translating Natural Language to SQL using Pointer-Generator Networks and How Decoding Order Matters 13 Nov 2018 · 0 repositories · arXiv:1811.05303
-
Macquarie University at BioASQ 6b: Deep learning and deep reinforcement learning for query-based summarisation 1 Nov 2018 · 0 repositories
-
Learning Scheduling Algorithms for Data Processing Clusters 3 Oct 2018 · 2 repositories · arXiv:1810.01963Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
A Fourier View of REINFORCE 12 Aug 2018 · 0 repositories · arXiv:1808.03953
-
ARM: Augment-REINFORCE-Merge Gradient for Stochastic Binary Networks 30 Jul 2018 · 1 repository · arXiv:1807.11143
-
Learning Globally Optimized Object Detector via Policy Gradient 1 Jun 2018 · 0 repositories
-
Revisiting Reweighted Wake-Sleep for Models with Stochastic Control Flow 26 May 2018 · 1 repository · arXiv:1805.10469
-
ReGAN: RE[LAX|BAR|INFORCE] based Sequence Generation using GANs 8 May 2018 · 0 repositories · arXiv:1805.02788
-
Adversarial Training for Community Question Answer Selection Based on Multi-scale Matching 22 Apr 2018 · 0 repositories · arXiv:1804.08058
-
CoT: Cooperative Training for Generative Modeling of Discrete Data 11 Apr 2018 · 2 repositories · arXiv:1804.03782
-
Attention, Learn to Solve Routing Problems! 22 Mar 2018 · 15 repositories · arXiv:1803.08475Syntology official (archive's flag): 3 ran · 27 ran (of which 0 constructed an object rather than computing a result; 20 with no instrument failure: 3 honoured, 0 violated, 17 with no contract checked; 7 where Syntology's instrument failed) · 13 unverified (of 40 harvested samples) · 5 pointer-only (licence)
-
Action-dependent Control Variates for Policy Optimization via Stein Identity 1 Jan 2018 · 0 repositories
-
Adversarial Policy Gradient for Alternating Markov Games 1 Jan 2018 · 0 repositories
-
LEARNING TO ORGANIZE KNOWLEDGE WITH N-GRAM MACHINES 1 Jan 2018 · 0 repositories
-
Technical Report for E2E NLG Challenge 19 Dec 2017 · 0 repositories
-
Differentiable lower bound for expected BLEU score 13 Dec 2017 · 2 repositories · arXiv:1712.04708
-
Action-depedent Control Variates for Policy Optimization via Stein's Identity 30 Oct 2017 · 2 repositories · arXiv:1710.11198Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Energy-efficient Amortized Inference with Cascaded Deep Classifiers 10 Oct 2017 · 0 repositories · arXiv:1710.03368
-
Rapid Probabilistic Interest Learning from Domain-Specific Pairwise Image Comparisons 19 Jun 2017 · 0 repositories · arXiv:1706.05850
-
Learning Hard Alignments with Variational Inference 16 May 2017 · 0 repositories · arXiv:1705.05524
-
Inferring and Executing Programs for Visual Reasoning 10 May 2017 · 5 repositories · arXiv:1705.03633Syntology official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Stein Variational Policy Gradient 7 Apr 2017 · 0 repositories · arXiv:1704.02399
-
REBAR: Low-variance, unbiased gradient estimates for discrete latent variable models 21 Mar 2017 · 3 repositories · arXiv:1703.07370
-
Neural Symbolic Machines: Learning Semantic Parsers on Freebase with Weak Supervision (Short Version) 4 Dec 2016 · 0 repositories · arXiv:1612.01197
-
Self-critical Sequence Training for Image Captioning 2 Dec 2016 · 31 repositories · arXiv:1612.00563Syntology 12 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 7 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples) · 3 pointer-only (licence)
-
Threshold Learning for Optimal Decision Making 1 Dec 2016 · 0 repositories
-
Improving Policy Gradient by Exploring Under-appreciated Rewards 28 Nov 2016 · 0 repositories · arXiv:1611.09321
-
Neural Symbolic Machines: Learning Semantic Parsers on Freebase with Weak Supervision 31 Oct 2016 · 2 repositories · arXiv:1611.00020
-
Episodic Exploration for Deep Deterministic Policies: An Application to StarCraft Micromanagement Tasks 10 Sep 2016 · 0 repositories · arXiv:1609.02993
-
End-to-end Learning of Action Detection from Frame Glimpses in Videos 22 Nov 2015 · 1 repository · arXiv:1511.06984
-
Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation 15 Aug 2013 · 2 repositories · arXiv:1308.3432
-
Analysis and Improvement of Policy Gradient Estimation 1 Dec 2011 · 0 repositories