{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2608-03796","title":"Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss","arxiv_id":"2608.03796","date":"2026-08-04","proceeding":null,"authors":["Bakbergen Ryskulov","Iker García-Ferrero","David Montero","David Jansen","Ali Hashemi","Jezabel R. Garcia","Antonio Tiene","Román Orús"],"abstract":"Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely trained from scratch: a compressed model is usually recovered through knowledge distillation (KD). This recovery step largely decides the final quality, yet it is expensive. We present a practitioner's study of how to make distillation training efficient, organised around two systems contributions. First, we show that offline KD (caching the teacher's top-$K$ logits once and training the student against the cache) matches online distillation at near-identical training loss while removing the teacher from memory, running about 29\\% faster per iteration, and reaching up to 41\\% higher throughput on a single H200 GPU. Second, we introduce a \\emph{fused, chunked KL loss} that never materialises the full vocabulary-sized logit tensor, making peak memory linear in the sequence length. This removes the memory spike that otherwise caps context length and lets us train at four times the context (32{,}768 tokens) on a single GPU. A separate output-head-only toy benchmark isolates the loss kernel and confirms its memory and iteration-rate scaling from 4K to 256K tokens. Together these make large-scale healing and hundreds of ablations affordable. We also report supporting ablations on loss design and sequence packing. We release our chunked-loss implementation: https://github.com/CompactifAI/Full-Chunked-KL-Loss.","url_abs":"https://arxiv.org/abs/2608.03796","url_pdf":"https://arxiv.org/pdf/2608.03796","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2608.03796","atlas_url":"https://app.syntology.ai/?focus=2608.03796","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2608.03796"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/CompactifAI/Full-Chunked-KL-Loss","reach":null}],"summary":{"ran":1,"ran_honours":2,"unverified":1},"by_repo_kind":{"found_in_text":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":4,"samples":[{"code_sha256_prefix":"dcebe0ca6417b0bd","entry":"_SparseKLHiddenStatesChunkedFunction","repo":"CompactifAI/Full-Chunked-KL-Loss","repo_kind":"found_in_text","path":"full_chunked.py","file_url":"https://github.com/CompactifAI/Full-Chunked-KL-Loss/blob/HEAD/full_chunked.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"dcebe0ca6417b0bd"}},{"code_sha256_prefix":"0d57dc1dc3c4e167","entry":"tensor_parallel_rank","repo":"CompactifAI/Full-Chunked-KL-Loss","repo_kind":"found_in_text","path":"full_chunked.py","file_url":"https://github.com/CompactifAI/Full-Chunked-KL-Loss/blob/HEAD/full_chunked.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0d57dc1dc3c4e167"}},{"code_sha256_prefix":"656a68efe430a488","entry":"tensor_parallel_size","repo":"CompactifAI/Full-Chunked-KL-Loss","repo_kind":"found_in_text","path":"full_chunked.py","file_url":"https://github.com/CompactifAI/Full-Chunked-KL-Loss/blob/HEAD/full_chunked.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"656a68efe430a488"}},{"code_sha256_prefix":"51c51b0e24ceb932","entry":"full_chunked_loss","repo":"CompactifAI/Full-Chunked-KL-Loss","repo_kind":"found_in_text","path":"full_chunked.py","file_url":"https://github.com/CompactifAI/Full-Chunked-KL-Loss/blob/HEAD/full_chunked.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"51c51b0e24ceb932"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":{"kind":"leaderboard_placements","source":"Syntology's leaderboard-shaped extractor over the paper's own arXiv-HTML tables: a model pointed at a cell, the number was read from that cell and checked against the board's metric, dataset, split and scale, and an independent check accepted the entry; not reviewed by the paper's authors or the archive's editors","extractor_model":"global.anthropic.claude-sonnet-4-5-20250929-v1:0","verifier_model":null,"prompt_sha":"fa63d4bb9d755694","coverage":{"sentence":"Syntology has checked 6,885 of the 9,623 papers on this site that are newer than the archive; results from the others appear after they are checked.","papers_newer_than_archive":9623,"papers_checked":6885},"entries":[],"not_placed":{"boards":0,"rejected_by_independent_check":0,"refused_by_a_rule":0,"check_did_not_answer":0,"proposed_without_a_cell":0,"declined_by_site":0}}}