Papers › LastOPD: Taming Collapse in Latent On-Policy Distillation

LastOPD: Taming Collapse in Latent On-Policy Distillation

23 Sep 2026arXiv:2609.28845added by Syntology

Jie Yang, Zhengyu Fang, Zelin Xu, Jiarui Sun, Xiran Fan, Junpeng Wang, Liang Wang, Qinghua Liu, Yiwei Cai, Yan Zheng

Title, abstract, authors and date from arXiv's metadata (CC0); this paper is not in the Papers with Code archive (frozen 2025-07-28).

On-policy distillation (OPD) corrects a student on the responses it writes, but its signal is the teacher's next-token distribution: it tells the student what the teacher says but misses how it thinks. Latent supervision promises the missing part by aligning the student's latent states to the teacher's. Recent methods such as OPRD bring this signal into on-policy distillation. However, we observe two failures of this recipe when distilling Qwen3-4B and Qwen3-8B into Qwen3-1.7B-Base. Early gain, late collapse: latent supervision alone lifts MATH-500 accuracy from 25 to 46 in 10 steps, but subsequent training degrades performance down to 11 with no recovery. Better alignment, worse behavior: although the alignment metric steadily improves throughout this collapse, the most aligned model turns out to be the worst performing. Further analysis suggests a mismatch in how the latent signal is applied: layers paired by depth play different roles in the two models, so continued alignment may pull the student toward teacher states it cannot understand. To address this, we propose LastOPD, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD. This keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in. Extensive experiments show that LastOPD improves MATH-500 over token-only OPD by 5.55 and 4.02 points with the 4B and 8B teachers, leads on most held-out datasets, and reaches the final score of token-only OPD in about half the steps. Code is available at https://github.com/Muyiiiii/LastOPD.

PaperPDF

In Syntology View this paper on Syntology, its page in Syntology's graph. That page lists the repositories linked to the paper, the abstract and the calls for agents.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, on Syntology's MCP service (how to connect):

Code

Muyiiiii/LastOPD found in paper text by SyntologySyntology: not harvested report

Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Syntology holds the repository link but has not harvested or run code from it.

Results from the paper

The Papers with Code archive ends with its 2025-07-28 snapshot. This paper's arXiv identifier, 2609.28845, was issued in September 2026, after that date, so the archive has no leaderboard rows for it.

Placed on leaderboards by Syntology Syntology

Syntology's extractor, the model Claude Sonnet 4.5, judged this paper's tables to report on 2 leaderboards (a model's judgement, not a result) and has not placed the paper on any of them: Mathematical Reasoning · AIME24 (refused by a rule: one configuration gave two values for one column); Mathematical Reasoning · AMC23 (refused by a rule: one configuration gave two values for one column). What is not shown.

Syntology has checked 8,886 of the 9,662 papers on this site that are newer than the archive (for 3,717 of them no archive leaderboard matched the paper's tables, so there was nothing further to check); 733 were read and have nothing to place (arXiv has no HTML version of the paper, or that version has no tables), 41 could not be read (the extractor's reply could not be parsed), and results from the other 2 appear after they are checked.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections