| On-policy Distillation with Verifiable Reward added by Syntology |
2026-08 (from id) |
LeapLabTHU/OPDVR/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| Mismatch Matters: On-Policy Distillation Beyond Token Agreement added by Syntology |
2026-08 (from id) |
yzc-666/TIDE/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| Toward Plasticity-Preserving KL Regularization for Capability Retention in LLM Reinforcement Learning added by Syntology |
2026-08 (from id) |
Lumina04/CoKL/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples added by Syntology |
2026-07 (from id) |
Hesse73/ARMOR/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning added by Syntology |
2026-06 (from id) |
jinyangwu/OPID/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
MIT (permissive) |
| OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning added by Syntology |
2026-06 (from id) |
pangpang-xuan/OPERA/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| ExTra: Exploratory Trajectory Optimization for Language Model Reinforcement Learning added by Syntology |
2026-06 (from id) |
allen4747/extra/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability added by Syntology |
2026-06 (from id) |
hp-luo/STARE/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs added by Syntology |
2026-06 (from id) |
Jiayi-Pan/TinyZero/verl/trainer/ppo/core_algos.py 45d08b3300b3da88 |
ran
|
Apache-2.0 (permissive) |
| SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models added by Syntology |
2026-06 (from id) |
Zhaolutuan/SAW/verl/trainer/ppo/core_algos.py 45d08b3300b3da88 |
ran
|
Apache-2.0 (permissive) |
| OPRD: On-Policy Representation Distillation added by Syntology |
2026-06 (from id) |
ShenzhiYang2000/OPRD/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation added by Syntology |
2026-06 (from id) |
YuYingLi0/FiRe-OPD/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards added by Syntology |
2026-05 (from id) |
Lumina04/CARE/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation added by Syntology |
2026-05 (from id) |
caiyuchen-ustc/EffOPD/EffOPD/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| SEIF: Self-Evolving Reinforcement Learning for Instruction Following added by Syntology |
2026-05 (from id) |
Rainier-rq1/SEIF/verl/trainer/core_algos.py 33a3cccef610855c |
ran
|
Apache-2.0 (permissive) |
| When Less is Enough: Efficient Inference via Collaborative Reasoning added by Syntology |
2026-05 (from id) |
fairytale9/llm_bottleneck/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment added by Syntology |
2026-04 (from id) |
XMUDeepLIT/HEAL/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| CoSearch: Joint Training of Reasoning and Document Ranking via Reinforcement Learning for Agentic Search added by Syntology |
2026-04 (from id) |
snap-research/CoSearch/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| Seeing Isn't Believing: Mitigating Belief Inertia via Active Intervention in Embodied Agents added by Syntology |
2026-04 (from id) |
WangHanLinHenry/EVU/verl-agent/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| Can LLMs Learn to Reason Robustly under Noisy Supervision? added by Syntology |
2026-04 (from id) |
ShenzhiYang2000/OLR/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| Experience is the Best Teacher: Motivating Effective Exploration in Reinforcement Learning for LLMs added by Syntology |
2026-03 (from id) |
sikelifei/HeRL/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time added by Syntology |
2026-03 (from id) |
Jasper-Yan/SCRL/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
MIT (permissive) |
| REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge added by Syntology |
2026-03 (from id) |
YasminZhang/REAL/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents added by Syntology |
2026-03 (from id) |
unimpor/T3/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| Entropy-Aware On-Policy Distillation of Language Models added by Syntology |
2026-03 (from id) |
WLS04/EOPD/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning added by Syntology |
2026-03 (from id) |
LuckyyySTA/GOLF/golf/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| SaFeR-ToolKit: Structured Reasoning via Virtual Tool Calling for Multimodal Safety added by Syntology |
2026-03 (from id) |
hiyouga/EasyR1/verl/trainer/core_algos.py 7633fd09bcf8972a |
unverified |
Apache-2.0 (permissive) |
| MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning added by Syntology |
2026-02 (from id) |
FlyTune/MASPO-RL/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| Beyond Static Pipelines: Learning Dynamic Workflows for Text-to-SQL added by Syntology |
2026-02 (from id) |
Satissss/SquRL/verl/trainer/ppo/core_algos.py 45d08b3300b3da88 |
ran
|
Apache-2.0 (permissive) |
| Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning added by Syntology |
2026-02 (from id) |
wdqqdw/Echo/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| Agent-Omit: Adaptive Context Omission for Efficient LLM Agents added by Syntology |
2026-02 (from id) |
usail-hkust/Agent-Omit/AgentOmit-RL/verl/agent_trainer/ppo/core_algos.py 45d08b3300b3da88 |
ran
|
no licence file found · pointer only |
| MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning added by Syntology |
2026-01 (from id) |
meituan/MemOCR/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| Agentic reinforcement learning empowers next-generation chemical language models for molecular design and synthesis added by Syntology |
2026-01 (from id) |
HowardLi1984/ChemCraft/chemcraft_rl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| RelayLLM: Efficient Reasoning via Collaborative Decoding added by Syntology |
2026-01 (from id) |
Chengsong-Huang/RelayLLM/RL_stage/verl/trainer/core_algos.py 7633fd09bcf8972a |
unverified |
no licence file found · pointer only |
| AgentOCR: Reimagining Agent History via Optical Self-Compression added by Syntology |
2026-01 (from id) |
langfengQ/AgentOCR/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| AT 2 PO: Agentic Turn-based Policy Optimization via Tree Search added by Syntology |
2026-01 (from id) |
zzfoutofspace/ATPO/ATPO/verl_atpo/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| Ophiuchus: Incentivizing Tool-augmented "Think with Images" for Joint Medical Segmentation, Understanding and Reasoning added by Syntology |
2025-12 (from id) |
SII-zyj/Ophiuchus/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization added by Syntology |
2025-12 (from id) |
ivanniu/NSPO/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
no licence file found · pointer only |
| Multi-Agent Tool-Integrated Policy Optimization added by Syntology |
2025-10 (from id) |
mzf666/MATPO/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| Scaling RL to Long Videos |
10 Jul 2025 |
NVlabs/Long-RL/verl/trainer/core_algos.py 7633fd09bcf8972a |
unverified |
Apache-2.0 (permissive) |
| Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training |
7 Jul 2025 |
zhhvvv/rft_vs_sft/EasyR1/verl/trainer/core_algos.py 33a3cccef610855c |
ran
|
no licence file found · pointer only |
| SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models |
15 Jun 2025 |
xid32/SoundMind/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
MIT (permissive) |
| Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency |
2 Jun 2025 |
appletea233/temporal-r1/verl/trainer/core_algos.py d61a07d67d82ec38 |
unverified |
Apache-2.0 (permissive) |
| ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay |
22 May 2025 |
dvlab-research/arpo/verl/trainer/core_algos.py 33a3cccef610855c |
ran
|
Apache-2.0 (permissive) |
| GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning |
16 May 2025 |
yueliu1999/guardreasoner-vl/train/EasyR1/verl/trainer/core_algos.py b5973cf2aefac349 |
unverified |
MIT (permissive) |
| TTRL: Test-Time Reinforcement Learning |
22 Apr 2025 |
prime-rl/ttrl/verl/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
MIT (permissive) |
| OpenVLThinker: An Early Exploration to Complex Vision-Language Reasoning via Iterative Self-Improvement |
21 Mar 2025 |
yihedeng9/openvlthinker/EasyR1/verl/trainer/core_algos.py 7633fd09bcf8972a |
unverified |
Apache-2.0 (permissive) |
| Guided Stream of Search: Learning to Better Search with Language Models via Optimal Path Guidance |
3 Oct 2024 |
snu-mllab/guided-rest/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| HybridFlow: A Flexible and Efficient RLHF Framework |
28 Sep 2024 |
du-nlp-lab/lengthreward/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| arXiv:openreview_v70fTOqer2 |
|
liziniu/KnapsackRL/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| arXiv:openreview_lbaBsu0CaY |
|
Tunanzzz/Meerkat-VL/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |
| arXiv:openreview_Y1PXB8HBV7 |
|
youwyouw/mirl-main/verl/trainer/core_algos.py 7633fd09bcf8972a |
unverified |
Apache-2.0 (permissive) |
| arXiv:openreview_Tkdgg5uqNK |
|
ZhijianZhou/Disppo/recipe/spin/core_algos.py 757d109d90121142 |
ran
|
Apache-2.0 (permissive) |