{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2608-24696","title":"On-policy Distillation with Verifiable Reward","arxiv_id":"2608.24696","date":"2026-08-25","proceeding":null,"authors":["Wenze Lin","Jiale Zhao","Xitai Jiang","Songde Rao","Yining Li","Shenzhi Wang","Bingxiang He","Gao Huang"],"abstract":"Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but ignores trajectory correctness, limiting its performance to that of the teacher. Combining them is a promising direction: OPD supplies dense supervisory signals, while RLVR provides task-level correctness. Nevertheless, existing integrations often rely on weighted combination or heuristic switching, introducing extra hyperparameters and trade-offs. We propose On-policy Distillation with Verifiable Reward (OPDVR), a simple yet effective method that seamlessly combines OPD and RLVR without adding any hyperparameters. We first reformulate the implicit reward of sampled-token OPD based on trajectory correctness, then apply a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards---thereby aligning the distillation signal with task success while preserving the teacher's distributional guidance. Furthermore, our modification transforms sampled-token OPD into a proper RLVR method, making it readily combinable with any policy gradient algorithm, such as GRPO. Experiments on six reasoning benchmarks show that OPDVR consistently outperforms standard OPD. Our code is available at https://github.com/LeapLabTHU/OPDVR.","url_abs":"https://arxiv.org/abs/2608.24696","url_pdf":"https://arxiv.org/pdf/2608.24696","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2608.24696","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2608.24696"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/LeapLabTHU/OPDVR","reach":null}],"summary":{"ran":9,"ran_fixture":1},"by_repo_kind":{"found_in_text":{"samples":10,"ran":10,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":10,"samples":[{"code_sha256_prefix":"f220595adfc5e83c","entry":"compute_ce_dpo_loss_rm","repo":"LeapLabTHU/OPDVR","repo_kind":"found_in_text","path":"verl/recipe/prime/prime_core_algos.py","file_url":"https://github.com/LeapLabTHU/OPDVR/blob/HEAD/verl/recipe/prime/prime_core_algos.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f220595adfc5e83c"}},{"code_sha256_prefix":"4d91e4383716e42c","entry":"compute_detach_dpo_loss_rm","repo":"LeapLabTHU/OPDVR","repo_kind":"found_in_text","path":"verl/recipe/prime/prime_core_algos.py","file_url":"https://github.com/LeapLabTHU/OPDVR/blob/HEAD/verl/recipe/prime/prime_core_algos.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4d91e4383716e42c"}},{"code_sha256_prefix":"f19825983a523003","entry":"compute_online_dpo_loss","repo":"LeapLabTHU/OPDVR","repo_kind":"found_in_text","path":"verl/recipe/spin/core_algos.py","file_url":"https://github.com/LeapLabTHU/OPDVR/blob/HEAD/verl/recipe/spin/core_algos.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f19825983a523003"}},{"code_sha256_prefix":"41b6e81dc3ecba79","entry":"compute_onlinedpo_pref","repo":"LeapLabTHU/OPDVR","repo_kind":"found_in_text","path":"verl/recipe/spin/core_algos.py","file_url":"https://github.com/LeapLabTHU/OPDVR/blob/HEAD/verl/recipe/spin/core_algos.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"41b6e81dc3ecba79"}},{"code_sha256_prefix":"05a86ccb0b3876dd","entry":"compute_score_data_source","repo":"LeapLabTHU/OPDVR","repo_kind":"found_in_text","path":"verl/recipe/open_math_reasoning/compute_score.py","file_url":"https://github.com/LeapLabTHU/OPDVR/blob/HEAD/verl/recipe/open_math_reasoning/compute_score.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"05a86ccb0b3876dd"}},{"code_sha256_prefix":"8896f4e2bafe8abc","entry":"generate_config_from_args","repo":"LeapLabTHU/OPDVR","repo_kind":"found_in_text","path":"verl/verl/model_merger/base_model_merger.py","file_url":"https://github.com/LeapLabTHU/OPDVR/blob/HEAD/verl/verl/model_merger/base_model_merger.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8896f4e2bafe8abc"}},{"code_sha256_prefix":"94209884dcb2308e","entry":"get_dynamic_pipeline_shards","repo":"LeapLabTHU/OPDVR","repo_kind":"found_in_text","path":"verl/verl/model_merger/megatron_model_merger.py","file_url":"https://github.com/LeapLabTHU/OPDVR/blob/HEAD/verl/verl/model_merger/megatron_model_merger.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"94209884dcb2308e"}},{"code_sha256_prefix":"757d109d90121142","entry":"get_kl_controller","repo":"LeapLabTHU/OPDVR","repo_kind":"found_in_text","path":"verl/recipe/spin/core_algos.py","file_url":"https://github.com/LeapLabTHU/OPDVR/blob/HEAD/verl/recipe/spin/core_algos.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"757d109d90121142"}},{"code_sha256_prefix":"11c8e99168eeef2e","entry":"normalize","repo":"LeapLabTHU/OPDVR","repo_kind":"found_in_text","path":"verl/recipe/gkd/megatron_kl_loss.py","file_url":"https://github.com/LeapLabTHU/OPDVR/blob/HEAD/verl/recipe/gkd/megatron_kl_loss.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"11c8e99168eeef2e"}},{"code_sha256_prefix":"8a84f8cc921357e2","entry":"reward_func","repo":"LeapLabTHU/OPDVR","repo_kind":"found_in_text","path":"verl/recipe/r1/reward_score.py","file_url":"https://github.com/LeapLabTHU/OPDVR/blob/HEAD/verl/recipe/r1/reward_score.py","link_basis":"plan_row","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"DEP_MISSING","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8a84f8cc921357e2"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}