| On-policy Distillation with Verifiable Reward added by Syntology |
2026-08 (from id) |
LeapLabTHU/OPDVR/verl/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Mismatch Matters: On-Policy Distillation Beyond Token Agreement added by Syntology |
2026-08 (from id) |
yzc-666/TIDE/verl/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning added by Syntology |
2026-08 (from id) |
Lumina04/CoKL/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples added by Syntology |
2026-07 (from id) |
Hesse73/ARMOR/verl/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
no licence file found · pointer only |
| OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning added by Syntology |
2026-06 (from id) |
jinyangwu/OPID/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
MIT (permissive) |
| OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning added by Syntology |
2026-06 (from id) |
pangpang-xuan/OPERA/verl/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
no licence file found · pointer only |
| ExTra: Exploratory Trajectory Optimization for Language Model Reinforcement Learning added by Syntology |
2026-06 (from id) |
allen4747/extra/verl/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability added by Syntology |
2026-06 (from id) |
hp-luo/STARE/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| OPRD: On-Policy Representation Distillation added by Syntology |
2026-06 (from id) |
ShenzhiYang2000/OPRD/verl/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning added by Syntology |
2026-06 (from id) |
THUAIS-Lab/CHERRL/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation added by Syntology |
2026-06 (from id) |
YuYingLi0/FiRe-OPD/verl/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards added by Syntology |
2026-05 (from id) |
Lumina04/CARE/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation added by Syntology |
2026-05 (from id) |
caiyuchen-ustc/EffOPD/EffOPD/verl/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| When Less is Enough: Efficient Inference via Collaborative Reasoning added by Syntology |
2026-05 (from id) |
fairytale9/llm_bottleneck/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment added by Syntology |
2026-04 (from id) |
XMUDeepLIT/HEAL/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| CoSearch: Joint Training of Reasoning and Document Ranking via Reinforcement Learning for Agentic Search added by Syntology |
2026-04 (from id) |
snap-research/CoSearch/verl/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| Seeing Isn't Believing: Mitigating Belief Inertia via Active Intervention in Embodied Agents added by Syntology |
2026-04 (from id) |
WangHanLinHenry/EVU/verl-agent/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
no licence file found · pointer only |
| GIANTS: Generative Insight Anticipation from Scientific Literature added by Syntology |
2026-04 (from id) |
joyheyueya/giants/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| Can LLMs Learn to Reason Robustly under Noisy Supervision? added by Syntology |
2026-04 (from id) |
ShenzhiYang2000/OLR/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Experience is the Best Teacher: Motivating Effective Exploration in Reinforcement Learning for LLMs added by Syntology |
2026-03 (from id) |
sikelifei/HeRL/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time added by Syntology |
2026-03 (from id) |
Jasper-Yan/SCRL/verl/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
MIT (permissive) |
| REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge added by Syntology |
2026-03 (from id) |
YasminZhang/REAL/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents added by Syntology |
2026-03 (from id) |
unimpor/T3/verl/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
no licence file found · pointer only |
| Entropy-Aware On-Policy Distillation of Language Models added by Syntology |
2026-03 (from id) |
WLS04/EOPD/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning added by Syntology |
2026-03 (from id) |
LuckyyySTA/GOLF/golf/verl/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning added by Syntology |
2026-02 (from id) |
FlyTune/MASPO-RL/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning added by Syntology |
2026-02 (from id) |
wdqqdw/Echo/verl/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
no licence file found · pointer only |
| MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning added by Syntology |
2026-01 (from id) |
meituan/MemOCR/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended Generation added by Syntology |
2026-01 (from id) |
YJiangcm/RLVRR/openrlhf/reward_tools/reward_fn.py 6e9c69dc6dfeb4ae |
ran · honoured contract
|
Apache-2.0 (permissive) |
| Agentic reinforcement learning empowers next-generation chemical language models for molecular design and synthesis added by Syntology |
2026-01 (from id) |
HowardLi1984/ChemCraft/chemcraft_rl/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
no licence file found · pointer only |
| AgentOCR: Reimagining Agent History via Optical Self-Compression added by Syntology |
2026-01 (from id) |
langfengQ/AgentOCR/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| AT 2 PO: Agentic Turn-based Policy Optimization via Tree Search added by Syntology |
2026-01 (from id) |
zzfoutofspace/ATPO/ATPO/verl_atpo/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
no licence file found · pointer only |
| Ophiuchus: Incentivizing Tool-augmented "Think with Images" for Joint Medical Segmentation, Understanding and Reasoning added by Syntology |
2025-12 (from id) |
SII-zyj/Ophiuchus/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization added by Syntology |
2025-12 (from id) |
ivanniu/NSPO/verl/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
no licence file found · pointer only |
| Multi-Agent Tool-Integrated Policy Optimization added by Syntology |
2025-10 (from id) |
mzf666/MATPO/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking added by Syntology |
2025-09 (from id) |
bytedance/EvoQuality/verl/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models |
15 Jun 2025 |
xid32/SoundMind/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
MIT (permissive) |
| RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling |
10 Jun 2025 |
bigai-nlco/rulereasoner/verl/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
MIT (permissive) |
| REARANK: Reasoning Re-ranking Agent via Reinforcement Learning |
26 May 2025 |
lezhang7/rearank/verl/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning |
20 May 2025 |
visual-agent/deepeyes/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| TTRL: Test-Time Reinforcement Learning |
22 Apr 2025 |
prime-rl/ttrl/verl/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
MIT (permissive) |
| Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models? |
2 Apr 2025 |
bigai-ai/ToM-RL/verl/utils/reward_score/explore_tom_2.py a5c20c9ff16e297d |
unverified |
MIT (permissive) |
| Guided Stream of Search: Learning to Better Search with Language Models via Optimal Path Guidance |
3 Oct 2024 |
snu-mllab/guided-rest/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| HybridFlow: A Flexible and Efficient RLHF Framework |
28 Sep 2024 |
du-nlp-lab/lengthreward/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| Model-based micro-data reinforcement learning: what are the crucial model properties and which model to choose? |
24 Jul 2021 |
ramp-kits/rl_simulator/benchmark/acrobot/reward_function.py d5e9182e08dfe611 |
unverified |
BSD-3-Clause recorded; this copy not marked cleared · pointer only |
| Energy-Based Imitation Learning |
20 Apr 2020 |
apexrl/EBIL-torch/rlkit/torch/ebil/ebil.py f97c59ae63058414 |
unverified |
MIT (permissive) |
| arXiv:openreview_v70fTOqer2 |
|
liziniu/KnapsackRL/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |
| arXiv:openreview_lbaBsu0CaY |
|
Tunanzzz/Meerkat-VL/recipe/r1/reward_score.py 8a84f8cc921357e2 |
ran · fixture could not drive it
|
Apache-2.0 (permissive) |
| arXiv:openreview_Tkdgg5uqNK |
|
ZhijianZhou/Disppo/recipe/r1/reward_score.py c30b2a96103abd2c |
ran
|
Apache-2.0 (permissive) |