Methods › Reinforcement Learning › Off-Policy TD Control › Q-Learning › Papers where code ran, page 1
Q-Learning
Papers archive 2025-07-28
archive papers tagged: 1,734 · with a code link: 464 · where Syntology ran a sample: 126 (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (126 of 1,734 tagged: 105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 1 of 2: papers 1 to 100 of the 126 tagged papers where Syntology ran at least one harvested sample (105 with a run with no instrument failure, 21 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Meta-Black-Box-Optimization through Offline Q-function Learning 4 May 2025 · 1 repository · arXiv:2505.02010Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy 8 Feb 2025 · 1 repository · arXiv:2502.05450Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
Flow Q-Learning 4 Feb 2025 · 2 repositories · arXiv:2502.02538Syntology official (archive's flag): 1 ran · 5 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 5 pointer-only (licence)
-
Temporal-Difference Learning Using Distributed Error Signals 6 Nov 2024 · 1 repository · arXiv:2411.03604Syntology official (archive's flag): 2 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 14 unverified (of 24 harvested samples) · 1 pointer-only (licence)
-
Beyond The Rainbow: High Performance Deep Reinforcement Learning on a Desktop PC 6 Nov 2024 · 3 repositories · arXiv:2411.03820Syntology official (archive's flag): 3 ran · 16 ran (of which 10 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 3 where Syntology's instrument failed) · 9 unverified (of 25 harvested samples) · 21 pointer-only (licence)
-
CALE: Continuous Arcade Learning Environment 31 Oct 2024 · 1 repository · arXiv:2410.23810Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Q-Distribution guided Q-learning for offline reinforcement learning: Uncertainty penalized Q-value via consistency model 27 Oct 2024 · 1 repository · arXiv:2410.20312Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Streaming Deep Reinforcement Learning Finally Works 18 Oct 2024 · 1 repository · arXiv:2410.14606Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Efficient and Private Marginal Reconstruction with Local Non-Negativity 1 Oct 2024 · 1 repository · arXiv:2410.01091Syntology official (archive's flag): 17 ran · 17 ran (of which 2 constructed an object rather than computing a result; 12 with no instrument failure: 10 honoured, 0 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 8 unverified (of 25 harvested samples) · 25 pointer-only (licence)
-
A Multi-Step Minimax Q-learning Algorithm for Two-Player Zero-Sum Markov Games 5 Jul 2024 · 1 repository · arXiv:2407.04240Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Simplifying Deep Temporal Difference Learning 5 Jul 2024 · 1 repository · arXiv:2407.04811Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Q-Adapter: Customizing Pre-trained LLMs to New Preferences with Forgetting Mitigation 4 Jul 2024 · 1 repository · arXiv:2407.03856Syntology official (archive's flag): 2 ran · 13 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 1 pointer-only (licence)
-
Diffusion Policies creating a Trust Region for Offline Reinforcement Learning 30 May 2024 · 1 repository · arXiv:2405.19690Syntology official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 6 where Syntology's instrument failed) · 3 unverified (of 20 harvested samples) · 20 pointer-only (licence)
-
Scalable Online Exploration via Coverability 11 Mar 2024 · 1 repository · arXiv:2403.06571Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Efficient Episodic Memory Utilization of Cooperative Multi-Agent Reinforcement Learning 2 Mar 2024 · 1 repository · arXiv:2403.01112Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
Conservative and Risk-Aware Offline Multi-Agent Reinforcement Learning 13 Feb 2024 · 1 repository · arXiv:2402.08421Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Improving Token-Based World Models with Parallel Observation Prediction 8 Feb 2024 · 1 repository · arXiv:2402.05643Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Averaging n-step Returns Reduces Variance in Reinforcement Learning 6 Feb 2024 · 0 repositories · arXiv:2402.03903Syntology 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 6 pointer-only (licence)
-
Q-Star Meets Scalable Posterior Sampling: Bridging Theory and Practice via HyperAgent 5 Feb 2024 · 3 repositories · arXiv:2402.10228Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples)
-
SPQR: Controlling Q-ensemble Independence with Spiked Random Model for Reinforcement Learning 6 Jan 2024 · 1 repository · arXiv:2401.03137Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples)
-
AdsorbRL: Deep Multi-Objective Reinforcement Learning for Inverse Catalysts Design 4 Dec 2023 · 1 repository · arXiv:2312.02308Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Reinforcement Learning for Wildfire Mitigation in Simulated Disaster Environments 27 Nov 2023 · 1 repository · arXiv:2311.15925Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Selectively Sharing Experiences Improves Multi-Agent Reinforcement Learning 1 Nov 2023 · 1 repository · arXiv:2311.00865Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Weakly Coupled Deep Q-Networks 28 Oct 2023 · 0 repositories · arXiv:2310.18803Syntology 5 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Towards Robust Offline Reinforcement Learning under Diverse Data Corruption 19 Oct 2023 · 2 repositories · arXiv:2310.12955Syntology official (archive's flag): 2 ran · 5 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Boosting Continuous Control with Consistency Policy 10 Oct 2023 · 1 repository · arXiv:2310.06343Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Pre-training with Synthetic Data Helps Offline Reinforcement Learning 1 Oct 2023 · 1 repository · arXiv:2310.00771Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Counterfactual Conservative Q Learning for Offline Multi-agent Reinforcement Learning 22 Sep 2023 · 1 repository · arXiv:2309.12696Syntology official (archive's flag): 6 ran · 7 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Towards Few-shot Coordination: Revisiting Ad-hoc Teamplay Challenge In the Game of Hanabi 20 Aug 2023 · 1 repository · arXiv:2308.10284Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Robust Multi-Agent Reinforcement Learning with State Uncertainty 30 Jul 2023 · 1 repository · arXiv:2307.16212Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
When should we prefer Decision Transformers for Offline Reinforcement Learning? 23 May 2023 · 1 repository · arXiv:2305.14550Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
IDQL: Implicit Q-Learning as an Actor-Critic Method with Diffusion Policies 20 Apr 2023 · 1 repository · arXiv:2304.10573Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
H-TSP: Hierarchically Solving the Large-Scale Travelling Salesman Problem 19 Apr 2023 · 1 repository · arXiv:2304.09395Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Modeling Moral Choices in Social Dilemmas with Multi-Agent Reinforcement Learning 20 Jan 2023 · 2 repositories · arXiv:2301.08491Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Extreme Q-Learning: MaxEnt RL without Entropy 5 Jan 2023 · 4 repositories · arXiv:2301.02328Syntology 8 ran (of which 5 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples) · 7 pointer-only (licence)
-
A Machine with Short-Term, Episodic, and Semantic Memory Systems 5 Dec 2022 · 1 repository · arXiv:2212.02098Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
Welfare and Fairness in Multi-objective Reinforcement Learning 30 Nov 2022 · 1 repository · arXiv:2212.01382Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
ACE: Cooperative Multi-agent Q-learning with Bidirectional Action-Dependency 29 Nov 2022 · 1 repository · arXiv:2211.16068Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Applying Deep Reinforcement Learning to the HP Model for Protein Structure Prediction 27 Nov 2022 · 1 repository · arXiv:2211.14939Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
Solving Continuous Control via Q-learning 22 Oct 2022 · 1 repository · arXiv:2210.12566Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient 13 Oct 2022 · 1 repository · arXiv:2210.06718Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Towards Safe Mechanical Ventilation Treatment Using Deep Offline Reinforcement Learning 5 Oct 2022 · 1 repository · arXiv:2210.02552Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Diffusion Policies as an Expressive Policy Class for Offline Reinforcement Learning 12 Aug 2022 · 3 repositories · arXiv:2208.06193Syntology official (archive's flag): 1 ran · 11 ran (of which 5 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 18 harvested samples) · 10 pointer-only (licence)
-
A Deep Reinforcement Learning Approach for Finding Non-Exploitable Strategies in Two-Player Atari Games 18 Jul 2022 · 2 repositories · arXiv:2207.08894Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples)
-
On the Learning and Learnability of Quasimetrics 30 Jun 2022 · 2 repositories · arXiv:2206.15478Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Mildly Conservative Q-Learning for Offline Reinforcement Learning 9 Jun 2022 · 3 repositories · arXiv:2206.04745Syntology official: no sample here; runs from other or unrecorded repositories · 13 ran (of which 6 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 1 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 16 harvested samples) · 11 pointer-only (licence)
-
Offline RL for Natural Language Generation with Implicit Language Q Learning 5 Jun 2022 · 2 repositories · arXiv:2206.11871Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Efficient Reward Poisoning Attacks on Online Deep Reinforcement Learning 30 May 2022 · 1 repository · arXiv:2205.14842Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
EPiDA: An Easy Plug-in Data Augmentation Framework for High Performance Text Classification 24 Apr 2022 · 1 repository · arXiv:2204.11205Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
A Statistical Analysis of Polyak-Ruppert Averaged Q-learning 29 Dec 2021 · 1 repository · arXiv:2112.14582Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Faster Deep Reinforcement Learning with Slower Online Network 10 Dec 2021 · 1 repository · arXiv:2112.05848Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
Dropout Q-Functions for Doubly Efficient Reinforcement Learning 5 Oct 2021 · 2 repositories · arXiv:2110.02034Syntology official (archive's flag): 1 ran · 6 ran (of which 3 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Uncertainty-Based Offline Reinforcement Learning with Diversified Q-Ensemble 4 Oct 2021 · 5 repositories · arXiv:2110.01548Syntology 13 ran (of which 11 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 8 unverified (of 21 harvested samples) · 6 pointer-only (licence)
-
On the Estimation Bias in Double Q-Learning 29 Sep 2021 · 1 repository · arXiv:2109.14419Syntology official (archive's flag): 4 ran · 4 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Learning cortical representations through perturbed and adversarial dreaming 9 Sep 2021 · 1 repository · arXiv:2109.04261Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Deep Reinforcement Learning at the Edge of the Statistical Precipice 30 Aug 2021 · 3 repositories · arXiv:2108.13264Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 4 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Backprop-Free Reinforcement Learning with Active Neural Generative Coding 10 Jul 2021 · 1 repository · arXiv:2107.07046Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Computational Benefits of Intermediate Rewards for Goal-Reaching Policy Learning 8 Jul 2021 · 1 repository · arXiv:2107.03961Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Ensemble and Auxiliary Tasks for Data-Efficient Deep Reinforcement Learning 5 Jul 2021 · 1 repository · arXiv:2107.01904Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Q-Learning Lagrange Policies for Multi-Action Restless Bandits 22 Jun 2021 · 1 repository · arXiv:2106.12024Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 17 harvested samples)
-
Vision-Language Navigation with Random Environmental Mixup 15 Jun 2021 · 1 repository · arXiv:2106.07876Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
Efficient (Soft) Q-Learning for Text Generation with Limited Good Data 14 Jun 2021 · 1 repository · arXiv:2106.07704Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
DouZero: Mastering DouDizhu with Self-Play Deep Reinforcement Learning 11 Jun 2021 · 1 repository · arXiv:2106.06135Syntology official (archive's flag): 2 ran · 3 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Adapting to Reward Progressivity via Spectral Reinforcement Learning 29 Apr 2021 · 1 repository · arXiv:2104.14138Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 2 harvested samples) · 2 pointer-only (licence)
-
Adaptive Rational Activations to Boost Deep Reinforcement Learning 18 Feb 2021 · 4 repositories · arXiv:2102.09407Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
DFAC Framework: Factorizing the Value Function via Quantile Mixture for Multi-Agent Distributional Q-Learning 16 Feb 2021 · 1 repository · arXiv:2102.07936Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Acting in Delayed Environments with Non-Stationary Markov Policies 28 Jan 2021 · 2 repositories · arXiv:2101.11992Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples)
-
Randomized Ensembled Double Q-Learning: Learning Fast Without a Model 15 Jan 2021 · 6 repositories · arXiv:2101.05982Syntology community repositories only · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 5 harvested samples) · 3 pointer-only (licence)
-
Reinforcement Learning with Latent Flow 6 Jan 2021 · 2 repositories · arXiv:2101.01857Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Learning Guidance Rewards with Trajectory-space Smoothing 23 Oct 2020 · 2 repositories · arXiv:2010.12718Syntology community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Q-learning with Language Model for Edit-based Unsupervised Summarization 9 Oct 2020 · 1 repository · arXiv:2010.04379Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
QPLEX: Duplex Dueling Multi-Agent Q-Learning 3 Aug 2020 · 6 repositories · arXiv:2008.01062Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
PC-PG: Policy Cover Directed Exploration for Provable Policy Gradient Learning 16 Jul 2020 · 1 repository · arXiv:2007.08459Syntology 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 6 pointer-only (licence)
-
Revisiting Fundamentals of Experience Replay 13 Jul 2020 · 2 repositories · arXiv:2007.06700Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
The Mean-Squared Error of Double Q-Learning 9 Jul 2020 · 1 repository · arXiv:2007.05034Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Gradient Temporal-Difference Learning with Regularized Corrections 1 Jul 2020 · 1 repository · arXiv:2007.00611Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Self-Imitation Learning via Generalized Lower Bound Q-learning 12 Jun 2020 · 0 repositories · arXiv:2006.07442Syntology 16 ran (of which 4 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 2 violated, 11 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 20 harvested samples) · 11 pointer-only (licence)
-
Conservative Q-Learning for Offline Reinforcement Learning 8 Jun 2020 · 18 repositories · arXiv:2006.04779Syntology community repositories only · 31 ran (of which 5 constructed an object rather than computing a result; 28 with no instrument failure: 2 honoured, 0 violated, 26 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 34 harvested samples) · 19 pointer-only (licence)
-
Local and Global Explanations of Agent Behavior: Integrating Strategy Summaries with Saliency Maps 18 May 2020 · 1 repository · arXiv:2005.08874Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 15 harvested samples)
-
Spatial Action Maps for Mobile Manipulation 20 Apr 2020 · 1 repository · arXiv:2004.09141Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
DisCor: Corrective Feedback in Reinforcement Learning via Distribution Correction 16 Mar 2020 · 4 repositories · arXiv:2003.07305Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Optimistic Exploration even with a Pessimistic Initialisation 26 Feb 2020 · 1 repository · arXiv:2002.12174Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Interpretable End-to-end Urban Autonomous Driving with Latent Deep Reinforcement Learning 23 Jan 2020 · 4 repositories · arXiv:2001.08726Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
SLM Lab: A Comprehensive Benchmark and Modular Software Framework for Reproducible Deep Reinforcement Learning 28 Dec 2019 · 1 repository · arXiv:1912.12482Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement Learning 27 Oct 2019 · 1 repository · arXiv:1910.12179Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
ZPD Teaching Strategies for Deep Reinforcement Learning from Demonstrations 26 Oct 2019 · 2 repositories · arXiv:1910.12154Syntology 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Meta-Q-Learning 30 Sep 2019 · 2 repositories · arXiv:1910.00125Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Learn How to Cook a New Recipe in a New House: Using Map Familiarization, Curriculum Learning, and Bandit Feedback to Learn Families of Text-Based Adventure Games 13 Aug 2019 · 1 repository · arXiv:1908.04777Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog 30 Jun 2019 · 1 repository · arXiv:1907.00456Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Towards Empathic Deep Q-Learning 26 Jun 2019 · 1 repository · arXiv:1906.10918Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Boosting Soft Actor-Critic: Emphasizing Recent Experience without Forgetting the Past 10 Jun 2019 · 3 repositories · arXiv:1906.04009Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Reinforcement Learning with Low-Complexity Liquid State Machines 4 Jun 2019 · 1 repository · arXiv:1906.01695Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction 3 Jun 2019 · 3 repositories · arXiv:1906.00949Syntology 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 5 harvested samples) · 5 pointer-only (licence)
-
Solving NP-Hard Problems on Graphs with Extended AlphaGo Zero 28 May 2019 · 2 repositories · arXiv:1905.11623Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Finite-Sample Analysis of Nonlinear Stochastic Approximation with Applications in Reinforcement Learning 27 May 2019 · 1 repository · arXiv:1905.11425Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Multi-Agent Deep Reinforcement Learning for Large-scale Traffic Signal Control 11 Mar 2019 · 1 repository · arXiv:1903.04527Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments 7 Mar 2019 · 3 repositories · arXiv:1903.03176Syntology official: harvested, nothing ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Diagnosing Bottlenecks in Deep Q-learning Algorithms 26 Feb 2019 · 1 repository · arXiv:1902.10250Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
Making Deep Q-learning methods robust to time discretization 28 Jan 2019 · 1 repository · arXiv:1901.09732Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Deep Reinforcement Learning for Imbalanced Classification 5 Jan 2019 · 3 repositories · arXiv:1901.01379Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)