{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2606-06021","title":"OPRD: On-Policy Representation Distillation","arxiv_id":"2606.06021","date":"2026-06-04","proceeding":null,"authors":["Shenzhi Yang","Guangcheng Zhu","Bowen Song","Haobo Wang","Mingxuan Xia","Xing Zheng","Yingfan Ma","Zhongqi Chen","Weiqiang Wang","Junbo Zhao","Gang Chen"],"abstract":"On-policy distillation (OPD) supervises the student exclusively in the output space by matching next-token distributions. This paradigm suffers from two limitations: (i) a high-variance gradient estimator whose signal-to-noise ratio collapses as the student approaches the teacher, and (ii) an LM-head information bottleneck that discards the teacher's intermediate hidden states. We propose On-Policy Representation Distillation (OPRD), the first method to lift on-policy distillation into the hidden-state space. OPRD aligns student and teacher representations across selected layers on the same on-policy rollouts, providing dense, deterministic, per-layer supervision while bypassing the LM head entirely. Theoretically, OPRD provides a deterministic per-sample gradient, removing the token-level estimation variance that plagues OPD, and exposes structural information that any output-space objective necessarily discards. Empirically, OPRD closes the student-teacher gap on competition mathematics benchmarks (AIME 2024, AIME 2025, and AIMO), where every output-space baseline plateaus below the teacher, while training 1.44x faster and using up to 54% less memory. We further extend OPRD to the cross-architecture setting via OPRD-Bridge. By exploiting the observation that heterogeneous models share a low-rank representational structure, we construct a frozen projector pair that aligns representations across arbitrary depth and width mismatches, shifting the alignment from the output space (which depends on a shared vocabulary) to the representation space. We validate OPRD-Bridge on both cross-architecture (Qwen3-4B -> Qwen3-1.7B-Base) and cross-tokenizer (Phi-4-mini-reasoning -> Qwen3-1.7B-Base) settings, demonstrating successful knowledge transfer even when the vocabulary-based alignment channel is unavailable. Code: https://github.com/ShenzhiYang2000/OPRD.","url_abs":"https://arxiv.org/abs/2606.06021","url_pdf":"https://arxiv.org/pdf/2606.06021","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2606.06021","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2606.06021"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/ShenzhiYang2000/OPRD","reach":{"status":"ok"}}],"summary":{"ran":8,"ran_fixture":1},"by_repo_kind":{"found_in_text":{"samples":9,"ran":9,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":9,"samples":[{"code_sha256_prefix":"f220595adfc5e83c","entry":"compute_ce_dpo_loss_rm","repo":"ShenzhiYang2000/OPRD","repo_kind":"found_in_text","path":"verl/recipe/prime/prime_core_algos.py","file_url":"https://github.com/ShenzhiYang2000/OPRD/blob/HEAD/verl/recipe/prime/prime_core_algos.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f220595adfc5e83c"}},{"code_sha256_prefix":"4d91e4383716e42c","entry":"compute_detach_dpo_loss_rm","repo":"ShenzhiYang2000/OPRD","repo_kind":"found_in_text","path":"verl/recipe/prime/prime_core_algos.py","file_url":"https://github.com/ShenzhiYang2000/OPRD/blob/HEAD/verl/recipe/prime/prime_core_algos.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4d91e4383716e42c"}},{"code_sha256_prefix":"f19825983a523003","entry":"compute_online_dpo_loss","repo":"ShenzhiYang2000/OPRD","repo_kind":"found_in_text","path":"verl/recipe/spin/core_algos.py","file_url":"https://github.com/ShenzhiYang2000/OPRD/blob/HEAD/verl/recipe/spin/core_algos.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f19825983a523003"}},{"code_sha256_prefix":"41b6e81dc3ecba79","entry":"compute_onlinedpo_pref","repo":"ShenzhiYang2000/OPRD","repo_kind":"found_in_text","path":"verl/recipe/spin/core_algos.py","file_url":"https://github.com/ShenzhiYang2000/OPRD/blob/HEAD/verl/recipe/spin/core_algos.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"41b6e81dc3ecba79"}},{"code_sha256_prefix":"05a86ccb0b3876dd","entry":"compute_score_data_source","repo":"ShenzhiYang2000/OPRD","repo_kind":"found_in_text","path":"verl/recipe/open_math_reasoning/compute_score.py","file_url":"https://github.com/ShenzhiYang2000/OPRD/blob/HEAD/verl/recipe/open_math_reasoning/compute_score.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"05a86ccb0b3876dd"}},{"code_sha256_prefix":"8896f4e2bafe8abc","entry":"generate_config_from_args","repo":"ShenzhiYang2000/OPRD","repo_kind":"found_in_text","path":"verl/verl/model_merger/base_model_merger.py","file_url":"https://github.com/ShenzhiYang2000/OPRD/blob/HEAD/verl/verl/model_merger/base_model_merger.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8896f4e2bafe8abc"}},{"code_sha256_prefix":"757d109d90121142","entry":"get_kl_controller","repo":"ShenzhiYang2000/OPRD","repo_kind":"found_in_text","path":"verl/recipe/spin/core_algos.py","file_url":"https://github.com/ShenzhiYang2000/OPRD/blob/HEAD/verl/recipe/spin/core_algos.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"757d109d90121142"}},{"code_sha256_prefix":"11c8e99168eeef2e","entry":"normalize","repo":"ShenzhiYang2000/OPRD","repo_kind":"found_in_text","path":"verl/recipe/gkd/megatron_kl_loss.py","file_url":"https://github.com/ShenzhiYang2000/OPRD/blob/HEAD/verl/recipe/gkd/megatron_kl_loss.py","link_basis":"plan_row","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"11c8e99168eeef2e"}},{"code_sha256_prefix":"8a84f8cc921357e2","entry":"reward_func","repo":"ShenzhiYang2000/OPRD","repo_kind":"found_in_text","path":"verl/recipe/r1/reward_score.py","file_url":"https://github.com/ShenzhiYang2000/OPRD/blob/HEAD/verl/recipe/r1/reward_score.py","link_basis":"plan_row","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"DEP_MISSING","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"8a84f8cc921357e2"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}