| Safe Continual Reinforcement Learning in Non-stationary Environments added by Syntology |
2026-04 (from id) |
MACS-Research-Lab/safe-crl/code/Safe-Policy-Optimization/safepo/common/buffer.py 185c4bf8b0d68d6c |
unverified |
no licence file found · pointer only |
| TREX: Trajectory Explanations for Multi-Objective Reinforcement Learning added by Syntology |
2026-03 (from id) |
dilina-r/trex_xmorl/trex_cluster.py 2dc44fc961aebd9a |
unverified |
no licence file found · pointer only |
| Reinforcement Learning Gradients as Vitamin for Online Finetuning Decision Transformers |
31 Oct 2024 |
kaiyan289/rl_as_vitamin_for_online_decision_transformers/data.py f434bf3e9ec07356 |
unverified |
no licence file found · pointer only |
| A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks |
29 Oct 2024 |
ml-jku/lram/src/buffers/buffer_utils.py 47f553906dc6d889 |
unverified |
MIT (permissive) |
| Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model Disentanglement |
15 Oct 2024 |
NJU-RL/Meta-DT/meta_dt/dataset.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| Retrieval-Augmented Decision Transformer: External Memory for In-context RL |
9 Oct 2024 |
ml-jku/RA-DT/src/buffers/buffer_utils.py 47f553906dc6d889 |
unverified |
MIT (permissive) |
| Decision Mamba: A Multi-Grained State Space Model with Self-Evolution Regularization for Offline RL |
8 Jun 2024 |
aopolin-lv/DecisionMamba/experiment-d4rl/experiment.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-Thought |
31 May 2024 |
identical code first harvested elsewhere 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Q-value Regularized Transformer for Offline Reinforcement Learning |
27 May 2024 |
charleshsc/qt/experiment.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
Apache-2.0 (permissive) |
| Is Mamba Compatible with Trajectory Optimization in Offline Reinforcement Learning? |
20 May 2024 |
AndssY/DeMa/gym/experiment.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Reinforced Sequential Decision-Making for Sepsis Treatment: The POSNEGDM Framework with Mortality Classifier and Transformer |
12 Mar 2024 |
dipeshtamboli/posnegdm-reinforced-sequential-decision-making-for-sepsis-treatment/utils.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning |
8 Feb 2024 |
woooodyy/llm-reverse-curriculum-rl/R3_math/src/utils.py 7ffa253a733f9bfc |
ran
fingerprinted |
no licence file found · pointer only |
| Critic-Guided Decision Transformer for Offline Reinforcement Learning |
21 Dec 2023 |
sharkwyf/cgdt/data.py f434bf3e9ec07356 |
unverified |
MIT (permissive) |
| Unleashing the Power of Pre-trained Language Models for Offline Reinforcement Learning |
31 Oct 2023 |
srzer/LaMo-2023/experiment-d4rl/experiment.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Learning to Modulate pre-trained Models in RL |
26 Jun 2023 |
ml-jku/l2m/src/buffers/buffer_utils.py 47f553906dc6d889 |
unverified |
MIT (permissive) |
| Future-conditioned Unsupervised Pretraining for Decision Transformer |
26 May 2023 |
fffffarmer/pdt/src/data.py 7ae6e29065b3cc4b |
unverified |
MIT (permissive) |
| Beyond Reward: Offline Preference-guided Policy Optimization |
25 May 2023 |
bkkgbkjb/oppo/oppo/human/gym/experiment.py 7a477d732661f79a |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Beyond Reward: Offline Preference-guided Policy Optimization |
25 May 2023 |
bkkgbkjb/oppo/oppo/scripted/gym/experiment.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
no licence file found · pointer only |
| When should we prefer Decision Transformers for Offline Reinforcement Learning? |
23 May 2023 |
prajjwal1/rl_paradigm/exorl/dataset.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Revisiting the Minimalist Approach to Offline Reinforcement Learning |
16 May 2023 |
adamjelley/efficientofflinerl/algorithms/cql.py d0c4a0aa6c8b23c1 |
unverified |
Apache-2.0 (permissive) |
| Revisiting the Minimalist Approach to Offline Reinforcement Learning |
16 May 2023 |
adamjelley/efficientofflinerl/algorithms/edac.py 04ad4ada1fd13ab2 |
unverified |
Apache-2.0 (permissive) |
| Revisiting the Minimalist Approach to Offline Reinforcement Learning |
16 May 2023 |
adamjelley/efficientofflinerl/algorithms/sac_n.py e258d5be5c961c7a |
unverified |
Apache-2.0 (permissive) |
| Merging Decision Transformers: Weight Averaging for Forming Multi-Task Policies |
14 Mar 2023 |
daniellawson9999/merging-decision-transformers/decision-transformer/experiment.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| How Crucial is Transformer in Decision Transformer? |
26 Nov 2022 |
identical code first harvested elsewhere 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| On the Effect of Pre-training for Transformer in Different Modality on Offline Reinforcement Learning |
17 Nov 2022 |
machelreid/can-wikipedia-help-offline-rl/code/eval_model.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| On the Effect of Pre-training for Transformer in Different Modality on Offline Reinforcement Learning |
17 Nov 2022 |
t46/pre-training-different-modality-offline-rl/can-wikipedia-help-offline-rl/experiment.py 0d4fc87c387d9018 |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Towards Understanding How Machines Can Learn Causal Overhypotheses |
16 Jun 2022 |
cannylab/casual_overhypotheses/models/decision-transformer/experiment.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| When does return-conditioned supervised learning work for offline reinforcement learning? |
2 Jun 2022 |
davidbrandfonbrener/rcsl-paper/decision-transformer/gym/experiment.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Efficient Reward Poisoning Attacks on Online Deep Reinforcement Learning |
30 May 2022 |
yinglunxu/reward_poisoning_attack_drl/src/all_class.py f047da47c4530c5f |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Can Wikipedia Help Offline Reinforcement Learning? |
28 Jan 2022 |
identical code first harvested elsewhere 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| Generalized Decision Transformer for Offline Hindsight Information Matching |
19 Nov 2021 |
identical code first harvested elsewhere 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
licence of this copy not recorded |
| StARformer: Transformer with State-Action-Reward Representations for Visual Reinforcement Learning |
12 Oct 2021 |
elicassion/StARformer/gym/experiment.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Decision Transformer: Reinforcement Learning via Sequence Modeling |
2 Jun 2021 |
HzcIrving/DecisionTransformer_StepbyStep/utils.py 0161b27cbe3cc58d |
ran · fixture could not drive it
fingerprinted |
MIT (permissive) |
| Enhancing SAT solvers with glue variable predictions |
6 Jul 2020 |
jesse-michael-han/neuro-cadical/python/rl_loop.py 658d370f8dacf5e3 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| A Closer Look at Invalid Action Masking in Policy Gradient Algorithms |
25 Jun 2020 |
vwxyzjn/invalid-action-masking/invalid_action_masking/ppo_10x10.py 43887c9edd549830 |
ran · fixture could not drive it
|
MIT (permissive) |
| Reward-Conditioned Policies |
31 Dec 2019 |
TrentBrick/RewardConditionedUDRL/control/agent.py d90f855adb55e3ce |
unverified |
MIT (permissive) |
| Importance Sampling Policy Evaluation with an Estimated Behavior Policy |
4 Jun 2018 |
LARG/regression-importance-sampling/roboschool-experiments/common.py 128b553ce11b9774 |
unverified |
MIT (permissive) |
| Meta-Gradient Reinforcement Learning |
24 May 2018 |
RobvanGastel/meta-rl-algorithms/algos/mg_a2c/buffer.py 85f38abd91775356 |
unverified |
MIT (permissive) |
| arXiv:aaai_35251 |
|
Teddy298/continualworld-ppo/continualworld/ppo/core.py 13141eda72cfce23 |
unverified |
MIT (permissive) |