{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2601-22156","title":"Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts","arxiv_id":"2601.22156","date":"2026-01-29","proceeding":null,"authors":["Yingfa Chen","Zhen Leng Thai","Zihan Zhou","Zhu Zhang","Xingyu Shen","Shuo Wang","Chaojun Xiao","Xu Han","Zhiyuan Liu"],"abstract":"Hybrid Transformer architectures, which combine softmax attention blocks and recurrent neural networks (RNNs), have shown a desirable performance-throughput tradeoff for long-context modeling, but their adoption and studies are hindered by the prohibitive cost of large-scale pre-training from scratch. Some recent studies have shown that pre-trained softmax attention blocks can be converted into RNN blocks through parameter transfer and knowledge distillation. However, these transfer methods require substantial amounts of training data (more than 10B tokens), and the resulting hybrid models also exhibit poor long-context performance, which is the scenario where hybrid models enjoy significant inference speedups over Transformer-based models. In this paper, we present HALO (Hybrid Attention via Layer Optimization), a pipeline for distilling Transformer models into RNN-attention hybrid models. We then present HypeNet, a hybrid architecture with superior length generalization enabled by a novel position encoding scheme (named HyPE) and various architectural modifications. We convert the Qwen3 series into HypeNet using HALO, achieving performance comparable to the original Transformer models while enjoying superior long-context performance and efficiency. The conversion requires just 2.3B tokens, less than 0.01% of their pre-training data","url_abs":"https://arxiv.org/abs/2601.22156","url_pdf":"https://arxiv.org/pdf/2601.22156","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2601.22156","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2601.22156"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/THUNLP/hybrid-linear-attention","reach":null}],"summary":{"unverified":7},"by_repo_kind":{"found_in_text":{"samples":7,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":7,"samples":[{"code_sha256_prefix":"879536e9e1571ff0","entry":"average_tasks","repo":"THUNLP/hybrid-linear-attention","repo_kind":"found_in_text","path":"attn-layer-selection/layer_analysis.py","file_url":"https://github.com/THUNLP/hybrid-linear-attention/blob/HEAD/attn-layer-selection/layer_analysis.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"879536e9e1571ff0"}},{"code_sha256_prefix":"1e4ead010f060e79","entry":"collect_layers","repo":"THUNLP/hybrid-linear-attention","repo_kind":"found_in_text","path":"attn-layer-selection/layer_analysis.py","file_url":"https://github.com/THUNLP/hybrid-linear-attention/blob/HEAD/attn-layer-selection/layer_analysis.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1e4ead010f060e79"}},{"code_sha256_prefix":"30bf3f3b16aff093","entry":"cu_seqlens_collate_fn","repo":"THUNLP/hybrid-linear-attention","repo_kind":"found_in_text","path":"halo/preparation.py","file_url":"https://github.com/THUNLP/hybrid-linear-attention/blob/HEAD/halo/preparation.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"30bf3f3b16aff093"}},{"code_sha256_prefix":"217db10d5f6c7d5a","entry":"extract_scores","repo":"THUNLP/hybrid-linear-attention","repo_kind":"found_in_text","path":"attn-layer-selection/layer_analysis.py","file_url":"https://github.com/THUNLP/hybrid-linear-attention/blob/HEAD/attn-layer-selection/layer_analysis.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"217db10d5f6c7d5a"}},{"code_sha256_prefix":"cb746560d2b54ee7","entry":"get_num_params","repo":"THUNLP/hybrid-linear-attention","repo_kind":"found_in_text","path":"halo/utils.py","file_url":"https://github.com/THUNLP/hybrid-linear-attention/blob/HEAD/halo/utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cb746560d2b54ee7"}},{"code_sha256_prefix":"94fbcdef5106cb83","entry":"load_rnn_checkpoint","repo":"THUNLP/hybrid-linear-attention","repo_kind":"found_in_text","path":"attn-layer-selection/convert.py","file_url":"https://github.com/THUNLP/hybrid-linear-attention/blob/HEAD/attn-layer-selection/convert.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"94fbcdef5106cb83"}},{"code_sha256_prefix":"fe8fd3f394926f19","entry":"run_inference","repo":"THUNLP/hybrid-linear-attention","repo_kind":"found_in_text","path":"attn-layer-selection/convert.py","file_url":"https://github.com/THUNLP/hybrid-linear-attention/blob/HEAD/attn-layer-selection/convert.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"fe8fd3f394926f19"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}