| SPEAR: Distilling Domain-Adaptive Reasoning Skeletons via Sequential Symbolic Alignment in Reinforcement Learning added by Syntology |
2026-08 (from id) |
zhuochunli/SPEAR/reward.py 3f97780f21c0b793 |
unverified |
no licence file found · pointer only |
| MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators added by Syntology |
2026-08 (from id) |
zzzzzzzzjj/MAVEN/src/train/reward_funcs.py cb8ca9226da3a936 |
ran
|
Apache-2.0 (permissive) |
| PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems added by Syntology |
2026-07 (from id) |
JAVA-Jiang/PowerAtlas/poweratlas/reward.py df8361c06df60d8e |
ran · honoured contract
|
licence not identified · pointer only |
| Zero-source LLM Hallucination Detection with Human-like Criteria Probing added by Syntology |
2026-06 (from id) |
TRISKEL10N/HCPD/src/open_r1/rewards.py ce353d566f3f2c9f |
ran
|
Apache-2.0 (permissive) |
| VideoSEG-O3: A Multi-turn Reinforcement Learning Framework for Reasoning Video Object Segmentation added by Syntology |
2026-06 (from id) |
Dmmm1997/VideoSEG-O3/projects/open_r1/rewards.py 09e6ed2ece2264f7 |
ran
|
Apache-2.0 (permissive) |
| Evolutionary Task Discovery: Advancing Reasoning Frontiers via Skill Composition and Complexity Scaling added by Syntology |
2026-05 (from id) |
liqinye/EvoTD/src/reward_function.py d67cad27811abe22 |
unverified |
MIT (permissive) |
| Strategy-Aware Optimization Modeling with Reasoning LLMs added by Syntology |
2026-05 (from id) |
rachhhhing/SAGE/reward_func/batch_score.py c15d4169879b293c |
ran
fingerprinted |
no licence file found · pointer only |
| Strategy-Aware Optimization Modeling with Reasoning LLMs added by Syntology |
2026-05 (from id) |
rachhhhing/SAGE/reward_func/batch_score_no_tem.py ef9fa7347607c290 |
ran
fingerprinted |
no licence file found · pointer only |
| OmniSapiens: A Foundation Model for Social Behavior Processing via Heterogeneity-Aware Relative Policy Optimization added by Syntology |
2026-02 (from id) |
MIT-MI/human_behavior_atlas/training/rl/reward_function/human_behaviour.py fd2112df3f149ab0 |
unverified |
no licence file found · pointer only |
| OmniSapiens: A Foundation Model for Social Behavior Processing via Heterogeneity-Aware Relative Policy Optimization added by Syntology |
2026-02 (from id) |
MIT-MI/human_behavior_atlas/training/rl/reward_function/human_behaviour_harpo.py b55511b458d41522 |
unverified |
no licence file found · pointer only |
| GOPO: Policy Optimization using Ranked Rewards added by Syntology |
2026-02 (from id) |
friendshipkim/gopo/src/open_r1/rewards.py ce353d566f3f2c9f |
ran
|
Apache-2.0 (permissive) |
| Alignment-Aware Model Adaptation via Feedback-Guided Optimization added by Syntology |
2026-02 (from id) |
facebookresearch/TruthRL/training/open-r1/src/open_r1/rewards.py ce353d566f3f2c9f |
ran
|
no licence file found · pointer only |
| 2 RELATED WORK Reinforcement learning has emerged as a powerful paradigm for enhancing the reasoning abilities of LLMs added by Syntology |
2026-01 (from id) |
zywang0104/Video-KTR/src/r1-v/src/open_r1/grpo.py 26720b6f9ee7eba4 |
unverified |
no licence file found · pointer only |
| RPO:Reinforcement Fine-Tuning with Partial Reasoning Optimization added by Syntology |
2026-01 (from id) |
yhz5613813/RPO/src/open_r1/rewards.py 0a5c0b6fb53cd48c |
unverified |
MIT (permissive) |
| Weather-R1: Logically Consistent Reinforcement Fine-Tuning for Multimodal Reasoning in Meteorology added by Syntology |
2026-01 (from id) |
Marcowky/Weather-R1/src/weather_r1/weather_r1_reward.py 12486ec2a6bdbe17 |
ran · honoured contract
fingerprinted |
MPL-2.0 (copyleft) · pointer only |
| RelayLLM: Efficient Reasoning via Collaborative Decoding added by Syntology |
2026-01 (from id) |
Chengsong-Huang/RelayLLM/RL_stage/data_filter.py e8ef1d6648c2b294 |
unverified |
no licence file found · pointer only |
| GRACE: Reinforcement Learning for Grounded Response and Abstention under Contextual Evidence added by Syntology |
2026-01 (from id) |
YiboZhao624/Grace/src/custom_reward_v2.py ca32d84f17d98414 |
ran · honoured contract
fingerprinted |
no licence file found · pointer only |
| Vision-Language Reasoning for Geolocalization: A Reinforcement Learning Approach added by Syntology |
2026-01 (from id) |
aialt/geo-r/src/open-r1-multimodal/src/open_r1/grpo.py e89f2eec0a18d4c2 |
unverified |
Apache-2.0 (permissive) |
| You Need Reasoning to Learn Reasoning: The Limitations of Label-Free RL in Weak Base Models added by Syntology |
2025-11 (from id) |
BorealisAI/CuMa/src/open_r1/rewards.py ce353d566f3f2c9f |
ran
|
no licence file found · pointer only |
| EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT added by Syntology |
2025-10 (from id) |
InternRobotics/EgoThinker/EgoThinker-RFT/src/open_r1/grpo.py 26720b6f9ee7eba4 |
unverified |
no licence file found · pointer only |
| EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT added by Syntology |
2025-10 (from id) |
InternRobotics/EgoThinker/EgoThinker-RFT/src/open_r1/grpo_video.py 06eb76b246b185f7 |
unverified |
no licence file found · pointer only |
| Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering added by Syntology |
2025-10 (from id) |
om-ai-lab/VLM-R1/src/open-r1-multimodal/src/open_r1/grpo.py e89f2eec0a18d4c2 |
unverified |
Apache-2.0 (permissive) |
| GTA1: GUI Test-time Scaling Agent |
8 Jul 2025 |
yan98/gta1/src/grpo_grounding.py b6cc386477eb6792 |
ran · honoured contract
|
no licence file found · pointer only |
| arXiv:2507.02834 |
2025-07 (from id) |
huggingface/open-r1/src/open_r1/rewards.py ce353d566f3f2c9f |
ran
|
Apache-2.0 (permissive) |
| arXiv:2507.02834 |
2025-07 (from id) |
dhcode-cpp/X-R1/src/x_r1/benchmark.py 2627a2d9d4808062 |
unverified |
Apache-2.0 (permissive) |
| DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation |
25 Jun 2025 |
apple/ml-diffucoder/src/open_r1/rewards.py ce353d566f3f2c9f |
ran
|
no licence file found · pointer only |
| GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents |
21 May 2025 |
yuqi-zhou/gui-g1/src/open-r1-multimodal/src/open_r1/grpo.py e89f2eec0a18d4c2 |
unverified |
Apache-2.0 (permissive) |
| VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank |
20 May 2025 |
tianhewu/visualquality-r1/src/open-r1-multimodal/src/open_r1/grpo.py e89f2eec0a18d4c2 |
unverified |
Apache-2.0 (permissive) |
| VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning |
18 May 2025 |
qiwang98/videorft/src/r1-v/src/open_r1/grpo.py 26720b6f9ee7eba4 |
unverified |
Apache-2.0 (permissive) |
| DRA-GRPO: Exploring Diversity-Aware Reward Adjustment for R1-Zero-Like Training of Large Language Models |
14 May 2025 |
xiwenc1/dra-grpo/src/open_r1/rewards.py 0a5c0b6fb53cd48c |
unverified |
MIT (permissive) |
| Tina: Tiny Reasoning Models via LoRA |
22 Apr 2025 |
shangshang-wang/tina/tina/post_train_hf/rewards.py 0a5c0b6fb53cd48c |
unverified |
Apache-2.0 (permissive) |
| SpaceR: Reinforcing MLLMs in Video Spatial Reasoning |
2 Apr 2025 |
ouyangkun10/spacer/SpaceR-SG-RLVR/src/r1-v/src/open_r1/SG-RLVR.py f23cbd84b69ec52c |
ran · honoured contract
|
licence not identified · pointer only |
| Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1 |
31 Mar 2025 |
tencentarc/seed-bench-r1/src/open_r1_egoplan/grpo.py 1fd4dec0e26848af |
unverified |
Apache-2.0 (permissive) |
| Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't |
20 Mar 2025 |
knoveleng/open-rs/src/open_r1/rewards.py 0a5c0b6fb53cd48c |
unverified |
MIT (permissive) |
| TimeZero: Temporal Video Grounding with Reasoning-Guided LVLM |
17 Mar 2025 |
www-ye/timezero/src/open_r1/grpo.py 26720b6f9ee7eba4 |
unverified |
Apache-2.0 (permissive) |
| Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering |
14 Mar 2025 |
xiaomi-research/r1-aqa/src/utils/rewards.py 54511e51a854f064 |
unverified |
Apache-2.0 (permissive) |
| MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark |
4 Sep 2024 |
opendatalab/pm4bench/src/pm4bench/qgo/reward.py c17d2901ecfbf4cb |
unverified |
Apache-2.0 (permissive) |
| Sequence-Augmented SE(3)-Flow Matching For Conditional Protein Backbone Generation |
30 May 2024 |
ding523/Curr_REFT/train_code/grpo/open-r1-multimodal/src/open_r1/Stage1_judge_math_resize.py baace3dc73e1da05 |
ran · honoured contract
|
no licence file found · pointer only |
| arXiv:2025.emnlp-main.67 |
|
isyuhaochen/RRC-DSCD/code/rl/rewards.py 70ba7655c5039f8f |
unverified |
MIT (permissive) |