| SCOPE-RL: Optimizing Reasoning Paths Before and After Success added by Syntology |
2026-07 (from id) |
tokencraft-lab/SCOPE-RL/verl/recipe/scope_rl/reward_score/step_quality.py e76cb5ca340e2e17 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation added by Syntology |
2026-06 (from id) |
YuYingLi0/FiRe-OPD/math_eval/eval_math.py 90b5c896e5eaea5e |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data added by Syntology |
2026-05 (from id) |
Celine-hxy/ATLAS/verl/verl/utils/reward_score/math_dapo.py 885dcaee42326ed4 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement added by Syntology |
2026-05 (from id) |
CuSO4-Chen/AKBE/AKBE/verl_akbe/verl/utils/reward_score/reward_em_betagrpo.py a459ca4d242e275d |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation added by Syntology |
2026-05 (from id) |
caiyuchen-ustc/EffOPD/EffOPD/math_eval/eval_math.py 90b5c896e5eaea5e |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| GraphPlanner: Graph Memory-Augmented Agentic Routing for Multi-Agent LLMs added by Syntology |
2026-04 (from id) |
ulab-uiuc/GraphPlanner/router_planner/shared/math_eval.py 9194e1f637041d86 |
ran
fingerprinted |
no licence file found · pointer only |
| TLoRA: Task-aware Low Rank Adaptation of Large Language Models added by Syntology |
2026-04 (from id) |
Rambo-Yi/TLora/evaluate/math/util.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| A Systematic Analysis of the Impact of Persona Steering on LLM Capabilities added by Syntology |
2026-04 (from id) |
cjia7/DPR/src/npti/eval/eval_math.py 0975c53e2a5fc2ab |
unverified |
no licence file found · pointer only |
| REAM: Merging Improves Pruning of Experts in LLMs added by Syntology |
2026-04 (from id) |
zai-org/glm-simple-evals/evals/math_eval.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| UpSkill: Mutual Information Skill Learning for Structured Response Diversity in LLMs added by Syntology |
2026-02 (from id) |
dshah02/upskill/src/DAPO_math_dapo.py 885dcaee42326ed4 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| UpSkill: Mutual Information Skill Learning for Structured Response Diversity in LLMs added by Syntology |
2026-02 (from id) |
dshah02/upskill/src/flex_extract.py 8e1ba71a728e180c |
unverified |
no licence file found · pointer only |
| IAPO: Information-Aware Policy Optimization for Token-Efficient Reasoning added by Syntology |
2026-02 (from id) |
YinhanHe123/IAPO/compute_reward.py 270295c5488675af |
unverified |
MIT (permissive) |
| ETS: Energy-Guided Test-Time Scaling for Training-Free RL Alignment added by Syntology |
2026-01 (from id) |
sheriyuo/ETS/aime24/utils.py 6d0bdd1050a5f9b9 |
unverified |
no licence file found · pointer only |
| DYNAACT: Large Language Model Reasoning with Dynamic Action Spaces added by Syntology |
2025-11 (from id) |
zhaoxlpku/DynaAct/math_utils.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Skywork-R1V3 Technical Report |
8 Jul 2025 |
seephys/seephys-project/vlmeval/api/bluelm_api.py a37f5e20fcdfc1ff |
unverified |
no licence file found · pointer only |
| SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning |
30 Jun 2025 |
spiral-rl/spiral/spiral/utils.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning |
20 May 2025 |
uw-nsl/tinyv/verl/verl/utils/reward_score/tinyv.py b6bc0ef47c33095e |
unverified |
MIT (permissive) |
| Rethinking the Role of Prompting Strategies in LLM Test-Time Scaling: A Perspective of Probability Theory |
16 May 2025 |
MraDonkey/rethinking_prompting/model.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Multi-Token Prediction Needs Registers |
15 May 2025 |
nasosger/mutor/language_modeling/src/eval/math_utils.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| arXiv:2503.05188 |
2025-03 (from id) |
BugMakerzzz/CRISP/crisp_reason.py fbe5c4948905ca15 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks |
2025-03 (from id) |
hengzzzhou/reso/ReSo/agent_graph/agent_graph.py 94659aac341e7312 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Big-Math: A Large-Scale, High-Quality Math Dataset for Reinforcement Learning in Language Models |
24 Feb 2025 |
synthlabsai/big-math/signals/rollouts_based_signals/math_eval.py abd7f82534e7b822 |
unverified |
MIT (permissive) |
| SIFT: Grounding LLM Reasoning in Contexts via Stickers |
19 Feb 2025 |
zhijie-group/sift/acc_stage2.py 90b5c896e5eaea5e |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities? |
17 Feb 2025 |
ZhiYuanZeng/test-time-scaling-eval/math_evaluator.py 90b5c896e5eaea5e |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| MasRouter: Learning to Route LLMs for Multi-Agent Systems |
16 Feb 2025 |
yanweiyue/masrouter/Datasets/math_dataset.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning |
10 Feb 2025 |
internlm/oreal/oreal/judgers/utils.py 90b5c896e5eaea5e |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| What Do Learning Dynamics Reveal About Generalization in LLM Reasoning? |
12 Nov 2024 |
katiekang1998/reasoning_generalization/math_eval_samples.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| InternLM2.5-StepProver: Advancing Automated Theorem Proving via Expert Iteration on Large-Scale LEAN Problems |
21 Oct 2024 |
internlm/internlm-math/agent/math_agent.py 90b5c896e5eaea5e |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Automatic Curriculum Expert Iteration for Reliable LLM Reasoning |
10 Oct 2024 |
salesforceairesearch/auto-cei/data/MATH/data_pre.py 8529d921bdf3f1aa |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning |
3 Oct 2024 |
identical code first harvested elsewhere 90b5c896e5eaea5e |
ran · our draft was wrong
fingerprinted |
licence of this copy not recorded |
| T2Vs Meet VLMs: A Scalable Multimodal Dataset for Visual Harmfulness Recognition |
29 Sep 2024 |
nctu-eva-lab/vhd11k/autogen/math_utils.py 735fc5841569ed55 |
ran
fingerprinted |
CC-BY-4.0 · pointer only |
| MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoning |
18 Sep 2024 |
dinobby/magicore/math_utils.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| BA-LoRA: Bias-Alleviating Low-Rank Adaptation to Mitigate Catastrophic Inheritance in Large Language Models |
8 Aug 2024 |
cyp-jlu-ai/ba-lora/inference/util.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback |
20 Jun 2024 |
kbsdjames/math-minos/evaluation/util.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B |
11 Jun 2024 |
trotsky1997/mathblackbox/run_with_earlystopping.py 90b5c896e5eaea5e |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| VB-LoRA: Extreme Parameter Efficient Fine-Tuning with Vector Banks |
24 May 2024 |
leo-yangli/VB-LoRA/math_instruction_tuning/instruction_tuning_eval/utils.py 90b5c896e5eaea5e |
ran · our draft was wrong
fingerprinted |
no licence file found · pointer only |
| Exploring the Compositional Deficiency of Large Language Models in Mathematical Reasoning |
5 May 2024 |
tongjingqi/MathTrap/eval/util.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Make Your LLM Fully Utilize the Context |
25 Apr 2024 |
microsoft/FILM/short_tasks/utils.py 37773130a24557b7 |
ran
fingerprinted |
MIT (permissive) |
| Advancing LLM Reasoning Generalists with Preference Trees |
2 Apr 2024 |
openbmb/eurus/eval/utils/util.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| NEFTune: Noisy Embeddings Improve Instruction Finetuning |
9 Oct 2023 |
akjindal53244/arithmo/eval/MATH/MATH_compute_metric_zero_shot_CoT.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Re-Reading Improves Reasoning in Large Language Models |
12 Sep 2023 |
tebmer/rereading-llm-reasoning/utils/util.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
Apache-2.0 (permissive) |
| Sparks of Artificial General Intelligence: Early experiments with GPT-4 |
22 Mar 2023 |
emrgnt-cmplxty/zero-shot-replication/zero_shot_replication/core/math_helpers.py a6e611557b1300a9 |
unverified |
Apache-2.0 (permissive) |
| Cost-Effective Hyperparameter Optimization for Large Language Model Generation Inference |
8 Mar 2023 |
kevin666aa/flaml/flaml/autogen/math_utils.py 735fc5841569ed55 |
ran
fingerprinted |
MIT recorded; this copy not marked cleared · pointer only |
| arXiv:openreview_QrC8OgQyOI |
|
viiika/Prism/Dream/Dream_Prism/metrics/gsmk8_eval.py a774929083345c9b |
unverified |
Apache-2.0 (permissive) |
| arXiv:2025.findings-emnlp.598 |
|
RUCAIBox/MMATH/utils.py 90b5c896e5eaea5e |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| arXiv:2025.emnlp-main.1660 |
|
dinobby/MAgICoRe/math_utils.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |
| arXiv:2024.findings-emnlp.407 |
|
amayuelas/multi-agent-attack/multiagent_debate/math_parsing.py 0b14c648516c38a7 |
ran · our draft was wrong
fingerprinted |
MIT (permissive) |