Browse State-of-the-Art › General Reinforcement Learning
General Reinforcement Learning
40 papers with code · 6 benchmarks · 7 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 40 papers with code (84 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 Dec 2017 62 repositories listed Syntology ran 13 of 17 samples · 4 unverified · 9 pointer-only (licence)The game of chess is the most widely-studied domain in the history of artificial intelligence.
-
26 Aug 2019 16 repositories listed Syntology ran 1 of 8 samples · 7 unverifiedOpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games.
-
13 Oct 2019 5 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Harnessing the transformer's ability to process long time horizons of information could provide a similar performance boost in partially observable reinforcement learning (RL) domains, but the large-scale transformers…
-
31 Aug 2018 5 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedDeveloping visual perception models for active agents and sensorimotor control are cumbersome to be done in the physical world, as existing algorithms are too slow to efficiently learn in real-time and robots are…
-
24 Nov 2017 5 repositories listedThis approach achieves a linear increase of the number of network outputs with the number of degrees of freedom by allowing a level of independence for each individual action dimension.
-
18 Feb 2021 4 repositories listed Syntology ran 6 of 6 samples · 0 unverified · 3 pointer-only (licence)Latest insights from biology show that intelligence not only emerges from the connections between neurons but that individual neurons shoulder more computational responsibility than previously anticipated.
-
21 Jun 2020 4 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedIn this work we aim to solve this problem by optimizing the efficiency and resource utilization of reinforcement learning algorithms instead of relying on distributed computation.
-
16 Oct 2023 3 repositories listed Syntology ran 8 of 13 samples · 5 unverified · 13 pointer-only (licence)ReMax can save about 46% GPU memory than PPO when training a 7B model and enables training on A800-80GB GPUs without the memory-saving offloading technique needed by PPO.
-
29 Oct 2022 2 repositories listedIn this paper, we present a Reinforcement Learning (RL) based methodology to DEtect and FIX (DeFIX) failures of an Imitation Learning (IL) agent by extracting infraction spots and re-constructing mini-scenarios on these…
-
15 Feb 2021 2 repositories listedSpatial memory, or the ability to remember and recall specific locations and objects, is central to autonomous agents' ability to carry out tasks in real environments.
-
7 Jul 2020 2 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedFor example, the common single-task sample-efficiency metric conflates improvements due to model-based learning with various other aspects, such as representation learning, making it difficult to assess true progress on…
-
10 Jun 2020 2 repositories listed Syntology ran 0 of 10 samples · 10 unverifiedThe challenge of developing powerful and general Reinforcement Learning (RL) agents has received increasing attention in recent years.
-
3 Jun 2019 2 repositories listedIn an effort to better understand the different ways in which the discount factor affects the optimization process in reinforcement learning, we designed a set of experiments to study each effect in isolation.
-
5 Mar 2019 2 repositories listedNumerous past works have tackled the problem of task-driven navigation.
-
4 Sep 2009 2 repositories listedThis paper introduces a principled approach for the design of a scalable general reinforcement learning agent.
-
21 May 2025 1 repository listedRecent advances such as DeepSeek R1-Zero highlight the effectiveness of incentive training, a reinforcement learning paradigm that computes rewards solely based on the final answer part of a language model's output,…
-
31 Mar 2025 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedWe propose Rec-R1, a general reinforcement learning framework that bridges large language models (LLMs) with recommendation systems through closed-loop optimization.
-
7 Nov 2024 1 repository listedIn this paper, a hypercube policy regularization framework is proposed, this method alleviates the constraints of policy constraint methods by allowing the agent to explore the actions corresponding to similar states in…
-
Kinetix: Investigating the Training of General Agents through Open-Ended Physics-Based Control Tasks30 Oct 2024 1 repository listed Syntology ran 24 of 33 samples · 9 unverifiedWhile large models trained with self-supervised learning on offline datasets have shown remarkable capabilities in text and image domains, achieving the same generalisation for agents that act in sequential decision…
-
4 Oct 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedRecently, it has been shown that it is possible to meta-learn update rules, with the hope of discovering algorithms that can perform well on a wide range of RL tasks.
-
6 Mar 2023 1 repository listed Syntology ran 0 of 13 samples · 13 unverifiedIn particular, we propose a general reinforcement learning-based backdoor attack framework where the attacker first trains a (non-myopic) attack policy using a simulator built upon its local data and common knowledge on…
-
17 Oct 2022 1 repository listedIn this paper, we propose a framework for intelligent vehicles to conduct JRC, with minimal prior knowledge of the system model and a tunable performance balance, in an environment where surrounding vehicles execute…
-
20 Jul 2022 1 repository listedWe test DMfD on a set of representative manipulation tasks for a 1-dimensional rope and a 2-dimensional cloth from the SoftGym suite of tasks, each with state and image observations.
-
31 Mar 2022 1 repository listedThe prevalent approach to unbiased click-based learning-to-rank (LTR) is based on counterfactual inverse-propensity-scoring (IPS) estimation.
-
14 Nov 2021 1 repository listedThe feasibility of making profitable trades on a single asset on stock exchanges based on patterns identification has long attracted researchers.
-
1 Sep 2021 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedIn this paper, we present IQ, i.
-
3 Jul 2021 1 repository listedIn this article we present the motivation and the core thesis towards the implementation of a Quantum Knowledge Seeking Agent (QKSA).
-
13 Feb 2021 1 repository listedWe present a novel interactive learning protocol that enables training request-fulfilling agents by verbally describing their activities.
-
28 Oct 2020 1 repository listed Syntology ran 0 of 5 samples · 5 unverifiedTo test this, we set forth the action hypergraph networks framework -- a class of functions for learning action representations in multi-dimensional discrete action spaces with a structural inductive bias.
-
29 Sep 2020 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedFor such complex tasks, the recently proposed RUDDER uses reward redistribution to leverage steps in the Q-function that are associated with accomplishing sub-tasks.
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections