Home › Code › get_generation_config

get_generation_config

Syntologyentry name in harvested coderead from the graph 2026-09-24

get_generation_config appears in the code Syntology harvested for 66 papers, as 12 distinct code bodies found in 66 places (a place is one code body under one paper). At least one of them ran in 1 of the papers; 0 of the code bodies carry a behaviour fingerprint.

What this page is not. Routines are grouped here by the exact string of their function or class name. Nothing asserts that two samples named get_generation_config do the same thing, share code, or are comparable; the name is a string, not an identity. Behaviour outputs (what a fingerprinted sample returned on the shared battery) are not in this export and are not shown here; the graph at syntology.ai holds them. "Ran" means executed on a synthesized fixture, not that the code is correct or reproduces a paper.

Samples Syntology

Syntology ran 1 of the 12 distinct code bodies named get_generation_config; 11 are unverified. One tile per status, in the site's fixed vocabulary, each code body counted once:

0ran · honoured contract
0ran · violated contract
0ran · our draft was wrong
0ran · fixture could not drive it
1ran
11unverified
0fingerprinted

Licence is a property of each copy, so it is counted per place: 18 of the 66 places are pointer only (Syntology does not serve that copy's text). This site shows no code text for any sample; every row below links to the file in its repository where the record names one.

“Ran” means the sample executed on a synthesized input; it does not mean the output is correct. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code, and those samples did run. The ran count above is every status except unverified, the same rule as each paper page.

Papers

66 papers shown of 66, newest first; 66 places in the table. A paper with no recorded date is placed by the month its arXiv id encodes, shown in the Date column as YYYY-MM (from id). One row per place: a paper whose repository defines the name more than once appears more than once, and the same code body held for several papers appears once under each, with the same status. Titles and dates are the archive's archive 2025-07-28 for papers in the archive, and the graph's for 49 papers added by Syntology; 2 papers have no page here and are shown by arXiv id only. Status and fingerprint are Syntology's record of each code body; licence is recorded for each place. The File cell ends with the code body's code_sha256, Syntology's identity for that exact code: an agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

PaperDateFileStatus SyntologyLicence
Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs added by Syntology 2026-09 (from id) hexixiang/MT-SDPO/training/verl/utils/model.py e6c72459395ae26a unverified Apache-2.0 (permissive)
Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents added by Syntology 2026-09 (from id) AlibabaResearch/SignalCoverageRL/verl/utils/model.py e4883a279c37b6cd unverified Apache-2.0 (permissive)
FARCA: Fact-Aligned Reliability-Aware Credit Assignment for Reinforcement Learning with Factual Supervision added by Syntology 2026-08 (from id) verl-project/verl/verl/utils/model.py 687521ab4cf0bf52 unverified Apache-2.0 (permissive)
Best Practice Critic Optimization BEST PRACTICE CRITIC OPTIMIZATION added by Syntology 2026-08 (from id) QPHutu/golden_critic/verl/utils/model.py 687521ab4cf0bf52 unverified Apache-2.0 (permissive)
Self-Improving Large Language Models via Progressive Experience Evolution added by Syntology 2026-08 (from id) rrrsj/SPEE/verl/utils/model.py 687521ab4cf0bf52 unverified Apache-2.0 (permissive)
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning added by Syntology 2026-07 (from id) xiuyilou/TACO/verl/utils/model.py e6c72459395ae26a unverified no licence file found · pointer only
DENSER ̸ = BETTER: LIMITS OF ON-POLICY SELF-DISTILLATION FOR CONTINUAL POST-TRAINING added by Syntology 2026-07 (from id) Moenupa/SDPO-CL/verl/utils/model.py e6c72459395ae26a unverified Apache-2.0 (permissive)
Stabilizing On-Policy Distillation for MLLM Reasoning with Global Normalization added by Syntology 2026-06 (from id) OPPO-Mente-Lab/GNDPO/verl/utils/model.py c9d58036da0286f2 unverified Apache-2.0 (permissive)
INFUSER: Influence-Guided Self-Evolution Improves Reasoning added by Syntology 2026-06 (from id) FFishy-git/INFUSER/verl/utils/model.py e6c72459395ae26a unverified MIT (permissive)
CollabBench: Benchmarking and Unleashing Collaborative Ability of LLMs with Diverse Players via Proactive Engagement added by Syntology 2026-06 (from id) langfengQ/verl-agent/verl/utils/model.py 5c059b45527bdff7 unverified Apache-2.0 (permissive)
Self-Distilled Policy Gradient added by Syntology 2026-06 (from id) lauyikfung/SDPG/verl/utils/model.py e6c72459395ae26a unverified Apache-2.0 (permissive)
Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification added by Syntology 2026-06 (from id) shanjf666/CoCoV/verl/verl/utils/model.py 68aa34a57b353cae unverified no licence file found · pointer only
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains added by Syntology 2026-05 (from id) ZiqiZhao1/ROSD/verl/utils/model.py e6c72459395ae26a unverified Apache-2.0 (permissive)
MERIT: Matching Expertise via Rubric-Informed Training for Reviewer Assignment added by Syntology 2026-05 (from id) Luli3220/MERIT/MERIT-Assessor/verl/utils/model.py e6c72459395ae26a unverified no licence file found · pointer only
Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning added by Syntology 2026-05 (from id) AlexFanw/LegalSearch-R1/verl/utils/model.py d3eec2b92fdb53f7 unverified Apache-2.0 (permissive)
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards added by Syntology 2026-05 (from id) Lumina04/CARE/verl/utils/model.py c9d58036da0286f2 unverified Apache-2.0 (permissive)
Hide to Guide: Learning via Semantic Masking added by Syntology 2026-05 (from id) mit-han-lab/SMEPO/verl/utils/model.py ea4b9efa02c4b431 unverified no licence file found · pointer only
Ranking-Aware Calibration for Reliable Multimodal Reinforcement Learning added by Syntology 2026-05 (from id) Stellaris167/RAC/verl/utils/model.py 687521ab4cf0bf52 unverified Apache-2.0 (permissive)
Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective added by Syntology 2026-05 (from id) horizon-llm/CTPO/verl/utils/model.py e6c72459395ae26a unverified Apache-2.0 (permissive)
One Refiner to Unlock Them All: Inference-Time Reasoning Elicitation via Reinforcement Query Refinement added by Syntology 2026-04 (from id) newera-xiao/ReQueR/verl/verl/utils/model.py d3eec2b92fdb53f7 unverified Apache-2.0 (permissive)
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning added by Syntology 2026-04 (from id) yuyongcan/DDRL/verl/verl/utils/model.py c9d58036da0286f2 unverified MIT (permissive)
Seeing Isn't Believing: Mitigating Belief Inertia via Active Intervention in Embodied Agents added by Syntology 2026-04 (from id) WangHanLinHenry/EVU/verl-agent/verl/utils/model.py 5c059b45527bdff7 unverified no licence file found · pointer only
GIANTS: Generative Insight Anticipation from Scientific Literature added by Syntology 2026-04 (from id) joyheyueya/giants/verl/utils/model.py 660713a8fa047736 unverified Apache-2.0 (permissive)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs? added by Syntology 2026-03 (from id) lasgroup/SDPO/verl/utils/model.py e6c72459395ae26a unverified Apache-2.0 (permissive)
TIPS: Turn-Level Information-Potential Reward Shaping for Search-Augmented LLMs added by Syntology 2026-03 (from id) ucsd-wang-lab-lm/tips/verl/utils/model.py d3eec2b92fdb53f7 unverified no licence file found · pointer only
What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time added by Syntology 2026-03 (from id) Jasper-Yan/SCRL/verl/verl/utils/model.py c9d58036da0286f2 unverified MIT (permissive)
Temporal Conflicts in LLMs: Reproducibility Insights from Unifying DYNAMICQA and MULAN added by Syntology 2026-03 (from id) terrierteam/temporal_conflicts/fact_mutability/bkp_inference.py 990a2f6afb605b7b unverified no licence file found · pointer only
Reinforcement Learning with Conditional Expectation Reward added by Syntology 2026-03 (from id) changyi7231/CER/verl/utils/model.py c9d58036da0286f2 unverified Apache-2.0 (permissive)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR added by Syntology 2026-03 (from id) Qwen-Applications/CLIPO/verl/utils/model.py 319697b76561a1e4 unverified no licence file found · pointer only
Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution added by Syntology 2026-03 (from id) ncbi-nlp/Med-V1/training/grpo/verl/utils/model.py 660713a8fa047736 unverified no licence file found · pointer only
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph added by Syntology 2026-02 (from id) Alibaba-NLP/VRAG/VRAG-RL/verl/utils/model.py 660713a8fa047736 unverified no licence file found · pointer only
Adaptive Reflection and Length Coordinated Penalty (ARLCP) STOP UNNECESSARY REFLECTION: TRAINING LRMS FOR EFFICIENT REASONING WITH ADAPTIVE REFLEC-TION AND LENGTH COORDINATED PENALTY added by Syntology 2026-02 (from id) ZeweiYu1/ARLCP/verl/utils/model.py 5c059b45527bdff7 unverified MIT (permissive)
Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning added by Syntology 2026-02 (from id) wdqqdw/Echo/verl/verl/utils/model.py 5c059b45527bdff7 unverified no licence file found · pointer only
Online Causal Kalman Filtering for Stable and Effective Policy Optimization added by Syntology 2026-02 (from id) shuohe1995/verl-kpo/verl/utils/model.py e6c72459395ae26a unverified Apache-2.0 (permissive)
Uncovering Cross-Objective Interference in Multi-Objective Alignment added by Syntology 2026-02 (from id) yining610/ctwa/verl/utils/model.py 319697b76561a1e4 unverified Apache-2.0 (permissive)
Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models added by Syntology 2026-02 (from id) Easy195/FaithRL/verl/utils/model.py 319697b76561a1e4 unverified no licence file found · pointer only
Rethinking the Trust Region in LLM Reinforcement Learning added by Syntology 2026-02 (from id) sail-sg/Stable-RL/verl/utils/model.py e6c72459395ae26a unverified Apache-2.0 (permissive)
PrAg-PO: Prompt Augmented Policy Optimization for Robust and Diverse Mathematical Reasoning added by Syntology 2026-02 (from id) wenquanlu/PrAg-PO/verl/utils/model.py 319697b76561a1e4 unverified no licence file found · pointer only
CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs added by Syntology 2026-02 (from id) Within-yao/CoBA-RL/verl/utils/model.py 319697b76561a1e4 unverified no licence file found · pointer only
CPMöbius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning added by Syntology 2026-02 (from id) thunlp/CPMobius/verl/utils/model.py 660713a8fa047736 unverified Apache-2.0 (permissive)
GRAPHDANCER: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training added by Syntology 2026-02 (from id) leopoldwhite/GraphDancer/verl/utils/model.py 319697b76561a1e4 unverified no licence file found · pointer only
AgentOCR: Reimagining Agent History via Optical Self-Compression added by Syntology 2026-01 (from id) langfengQ/AgentOCR/verl/utils/model.py 5c059b45527bdff7 unverified Apache-2.0 (permissive)
AT 2 PO: Agentic Turn-based Policy Optimization via Tree Search added by Syntology 2026-01 (from id) zzfoutofspace/ATPO/ATPO/verl_atpo/verl/utils/model.py 5c059b45527bdff7 unverified no licence file found · pointer only
Prioritized Replay for RL Post-training added by Syntology 2026-01 (from id) fatemi/verl/verl/utils/model.py e6c72459395ae26a unverified Apache-2.0 (permissive)
Meta-RL Induces Exploration in Language Agents added by Syntology 2025-12 (from id) mlbio-epfl/LaMer/verl/utils/model.py 5c059b45527bdff7 unverified no licence file found · pointer only
Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning added by Syntology 2025-11 (from id) aiming-lab/Agent0/Agent0-VL/verl/utils/model.py 660713a8fa047736 unverified Apache-2.0 (permissive)
Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents added by Syntology 2025-10 (from id) GuoqingWang1/IGPO/verl/utils/model.py 68aa34a57b353cae unverified Apache-2.0 (permissive)
Multi-Agent Tool-Integrated Policy Optimization added by Syntology 2025-10 (from id) mzf666/MATPO/verl/utils/model.py c9d58036da0286f2 unverified Apache-2.0 (permissive)
Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking added by Syntology 2025-09 (from id) bytedance/EvoQuality/verl/verl/utils/model.py d3eec2b92fdb53f7 unverified Apache-2.0 (permissive)
SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models 15 Jun 2025 xid32/SoundMind/verl/utils/model.py 5c059b45527bdff7 unverified MIT (permissive)
RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic Sampling 10 Jun 2025 bigai-nlco/rulereasoner/verl/verl/utils/model.py 68aa34a57b353cae unverified MIT (permissive)
ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL 30 May 2025 Franklin-Zhang0/ReasonGen-R1/verl/utils/model.py 660713a8fa047736 unverified Apache-2.0 (permissive)
REARANK: Reasoning Re-ranking Agent via Reinforcement Learning 26 May 2025 lezhang7/rearank/verl/verl/utils/model.py 68aa34a57b353cae unverified Apache-2.0 (permissive)
Behavior Injection: Preparing Language Models for Reinforcement Learning 25 May 2025 czp16/bridge-llm-reasoning/simple_verl/sverl/utils/model.py 660713a8fa047736 unverified MIT (permissive)
AdaCtrl: Towards Adaptive and Controllable Reasoning via Difficulty-Aware Budgeting 24 May 2025 joeying1019/adactrl/verl/utils/model.py 660713a8fa047736 unverified Apache-2.0 (permissive)
DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning 20 May 2025 visual-agent/deepeyes/verl/utils/model.py 68aa34a57b353cae unverified Apache-2.0 (permissive)
Reinforcement Learning for Reasoning in Large Language Models with One Training Example 29 Apr 2025 ypwang61/one-shot-rlvr/verl/utils/model.py 660713a8fa047736 unverified Apache-2.0 (permissive)
Efficient Reinforcement Finetuning via Adaptive Curriculum Learning 7 Apr 2025 uscnlp-lime/verl/verl/utils/model.py 660713a8fa047736 unverified Apache-2.0 (permissive)
CPPO: Accelerating the Training of Group Relative Policy Optimization-Based Reasoning Models 28 Mar 2025 lzhxmu/cppo/cppo_verl/verl/utils/model.py 660713a8fa047736 unverified Apache-2.0 (permissive)
Agent models: Internalizing Chain-of-Action Generation into Reasoning models 9 Mar 2025 adam-bjtu/autocoa/RL_training/verl/utils/model.py 660713a8fa047736 unverified MIT (permissive)
Self-rewarding correction for mathematical reasoning 26 Feb 2025 agentica-project/verl-pipeline/verl/utils/model.py 660713a8fa047736 unverified Apache-2.0 (permissive)
Flaming-hot Initiation with Regular Execution Sampling for Large Language Models 28 Oct 2024 yaof20/real/verl/utils/model.py 660713a8fa047736 unverified Apache-2.0 (permissive)
HybridFlow: A Flexible and Efficient RLHF Framework 28 Sep 2024 du-nlp-lab/lengthreward/verl/utils/model.py c9d58036da0286f2 unverified Apache-2.0 (permissive)
Democratizing Reasoning Ability: Tailored Learning from Large Language Model 20 Oct 2023 Raibows/Learn-to-Reason/utils.py 4b72086bac0f516f ran no licence file found · pointer only
arXiv:openreview_dzcDbh8ewp Xiao-Youth/LECTOR/verl/utils/model.py 5c059b45527bdff7 unverified Apache-2.0 (permissive)
arXiv:openreview_VZDlkXuIQc ShadeCloak/IP-GRM/verl/verl/utils/model.py 319697b76561a1e4 unverified Apache-2.0 (permissive)

This site shows no code text; each File cell links to the file on GitHub at the repository's current default branch, which may have changed since the harvest. "Pointer only" means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence cell for the reason. Per-sample records for a paper are on its paper page under "Code Syntology ran".

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections