{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/retentive-network-a-successor-to-transformer","title":"Retentive Network: A Successor to Transformer for Large Language Models","arxiv_id":"2307.08621","date":"2023-07-17","proceeding":null,"authors":["Yutao Sun","Li Dong","Shaohan Huang","Shuming Ma","Yuqing Xia","Jilong Xue","Jianyong Wang","Furu Wei"],"abstract":"In this work, we propose Retentive Network (RetNet) as a foundation architecture for large language models, simultaneously achieving training parallelism, low-cost inference, and good performance. We theoretically derive the connection between recurrence and attention. Then we propose the retention mechanism for sequence modeling, which supports three computation paradigms, i.e., parallel, recurrent, and chunkwise recurrent. Specifically, the parallel representation allows for training parallelism. The recurrent representation enables low-cost $O(1)$ inference, which improves decoding throughput, latency, and GPU memory without sacrificing performance. The chunkwise recurrent representation facilitates efficient long-sequence modeling with linear complexity, where each chunk is encoded parallelly while recurrently summarizing the chunks. Experimental results on language modeling show that RetNet achieves favorable scaling results, parallel training, low-cost deployment, and efficient inference. The intriguing properties make RetNet a strong successor to Transformer for large language models. Code will be available at https://aka.ms/retnet.","url_abs":"https://arxiv.org/abs/2307.08621v4","url_pdf":"https://arxiv.org/pdf/2307.08621v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"retentive-network-a-successor-to-transformer","repo_url":"https://github.com/microsoft/unilm","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"retentive-network-a-successor-to-transformer","repo_url":"https://github.com/Jamie-Stirling/RetNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"retentive-network-a-successor-to-transformer","repo_url":"https://github.com/catworldlee/gaussian-mixture-mask-attention","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"retentive-network-a-successor-to-transformer","repo_url":"https://github.com/fkodom/yet-another-retnet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"retentive-network-a-successor-to-transformer","repo_url":"https://github.com/lions-epfl/lion","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"retentive-network-a-successor-to-transformer","repo_url":"https://github.com/microsoft/torchscale","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"retentive-network-a-successor-to-transformer","repo_url":"https://github.com/osu-starlab/leapformer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"retentive-network-a-successor-to-transformer","repo_url":"https://github.com/sustcsonglin/flash-linear-attention","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"retentive-network-a-successor-to-transformer","repo_url":"https://github.com/prateekstark/retnet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":null,"task_name":"GPU"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2307.08621","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2307.08621"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/prateekstark/retnet","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/osu-starlab/leapformer","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/sustcsonglin/flash-linear-attention","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/fkodom/yet-another-retnet","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Jamie-Stirling/RetNet","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/microsoft/torchscale","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/catworldlee/gaussian-mixture-mask-attention","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lions-epfl/lion","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/microsoft/unilm","reach":null}],"summary":{"ran_honours":1,"ran":2,"unverified":2},"by_repo_kind":{"listed":{"samples":5,"ran":3,"repositories":3}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"94173bcb2aff5be6","entry":"On_attention_gaussian_mask","repo":"catworldlee/gaussian-mixture-mask-attention","repo_kind":"listed","path":"models/gmm.py","file_url":"https://github.com/catworldlee/gaussian-mixture-mask-attention/blob/HEAD/models/gmm.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"94173bcb2aff5be6"}},{"code_sha256_prefix":"03e99c761c545617","entry":"duplicate_interleave","repo":"Jamie-Stirling/RetNet","repo_kind":"listed","path":"src/xpos_relative_position.py","file_url":"https://github.com/Jamie-Stirling/RetNet/blob/HEAD/src/xpos_relative_position.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"03e99c761c545617"}},{"code_sha256_prefix":"e10a34a678e99e04","entry":"fixed_pos_embedding","repo":"Jamie-Stirling/RetNet","repo_kind":"listed","path":"src/xpos_relative_position.py","file_url":"https://github.com/Jamie-Stirling/RetNet/blob/HEAD/src/xpos_relative_position.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e10a34a678e99e04"}},{"code_sha256_prefix":"b2a514c2dffaca20","entry":"rotate_every_two","repo":"Jamie-Stirling/RetNet","repo_kind":"listed","path":"src/xpos_relative_position.py","file_url":"https://github.com/Jamie-Stirling/RetNet/blob/HEAD/src/xpos_relative_position.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b2a514c2dffaca20"}},{"code_sha256_prefix":"55b6a2a21f24cc60","entry":"transformer_1_3b","repo":"fkodom/yet-another-retnet","repo_kind":"listed","path":"scripts/benchmark_inference.py","file_url":"https://github.com/fkodom/yet-another-retnet/blob/HEAD/scripts/benchmark_inference.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"55b6a2a21f24cc60"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}