Methods › Reinforcement Learning › Policy Gradient Methods › REINFORCE › Papers where code ran, page 1
REINFORCE
Papers archive 2025-07-28
archive papers tagged: 185 · with a code link: 80 · where Syntology ran a sample: 24 (20 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (24 of 185 tagged: 20 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 1 of 1: papers 1 to 24 of the 24 tagged papers where Syntology ran at least one harvested sample (20 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Finding Visual Task Vectors 8 Apr 2024 · 1 repository · arXiv:2404.05729Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 9 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Re3val: Reinforced and Reranked Generative Retrieval 30 Jan 2024 · 0 repositories · arXiv:2401.16979Syntology 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models 16 Oct 2023 · 3 repositories · arXiv:2310.10505Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
ESRL: Efficient Sampling-based Reinforcement Learning for Sequence Generation 4 Aug 2023 · 2 repositories · arXiv:2308.02223Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples) · 3 pointer-only (licence)
-
Generalizing Few-Shot NAS with Gradient Matching 29 Mar 2022 · 1 repository · arXiv:2203.15207Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples)
-
Gradient Estimation with Discrete Stein Operators 19 Feb 2022 · 1 repository · arXiv:2202.09497Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Double Control Variates for Gradient Estimation in Discrete Latent Variable Models 9 Nov 2021 · 1 repository · arXiv:2111.05300Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
ReGen: Reinforcement Learning for Text and Knowledge Base Generation using Pretrained Language Models 27 Aug 2021 · 1 repository · arXiv:2108.12472Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
ARMS: Antithetic-REINFORCE-Multi-Sample Gradient for Binary Variables 28 May 2021 · 1 repository · arXiv:2105.14141Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
POMO: Policy Optimization with Multiple Optima for Reinforcement Learning 30 Oct 2020 · 3 repositories · arXiv:2010.16011Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
ConfuciuX: Autonomous Hardware Resource Assignment for DNN Accelerators using Reinforcement Learning 4 Sep 2020 · 1 repository · arXiv:2009.02010Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Angle-based Search Space Shrinking for Neural Architecture Search 28 Apr 2020 · 1 repository · arXiv:2004.13431Syntology 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Guided Dialog Policy Learning without Adversarial Learning in the Loop 7 Apr 2020 · 1 repository · arXiv:2004.03267Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
A Hybrid Stochastic Policy Gradient Algorithm for Reinforcement Learning 1 Mar 2020 · 1 repository · arXiv:2003.00430Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Neural Predictor for Neural Architecture Search 2 Dec 2019 · 2 repositories · arXiv:1912.00848Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
Exploiting Uncertainty of Loss Landscape for Stochastic Optimization 30 May 2019 · 1 repository · arXiv:1905.13200Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Interpretable Neural Predictions with Differentiable Binary Variables 20 May 2019 · 1 repository · arXiv:1905.08160Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
ARSM: Augment-REINFORCE-Swap-Merge Estimator for Gradient Backpropagation Through Categorical Variables 4 May 2019 · 1 repository · arXiv:1905.01413Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware 2 Dec 2018 · 23 repositories · arXiv:1812.00332Syntology official (archive's flag): 2 ran · 17 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 2 honoured, 0 violated, 12 with no contract checked; 3 where Syntology's instrument failed) · 10 unverified (of 27 harvested samples) · 4 pointer-only (licence)
-
Learning Scheduling Algorithms for Data Processing Clusters 3 Oct 2018 · 2 repositories · arXiv:1810.01963Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Attention, Learn to Solve Routing Problems! 22 Mar 2018 · 15 repositories · arXiv:1803.08475Syntology official (archive's flag): 3 ran · 27 ran (of which 0 constructed an object rather than computing a result; 20 with no instrument failure: 3 honoured, 0 violated, 17 with no contract checked; 7 where Syntology's instrument failed) · 13 unverified (of 40 harvested samples) · 5 pointer-only (licence)
-
Action-depedent Control Variates for Policy Optimization via Stein's Identity 30 Oct 2017 · 2 repositories · arXiv:1710.11198Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Inferring and Executing Programs for Visual Reasoning 10 May 2017 · 5 repositories · arXiv:1705.03633Syntology official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Self-critical Sequence Training for Image Captioning 2 Dec 2016 · 31 repositories · arXiv:1612.00563Syntology 12 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 7 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples) · 3 pointer-only (licence)