{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/factorized-attention-self-attention-with","title":"Efficient Attention: Attention with Linear Complexities","arxiv_id":"1812.01243","date":"2018-12-04","proceeding":null,"authors":["Zhuoran Shen","Mingyuan Zhang","Haiyu Zhao","Shuai Yi","Hongsheng Li"],"abstract":"Dot-product attention has wide applications in computer vision and natural language processing. However, its memory and computational costs grow quadratically with the input size. Such growth prohibits its application on high-resolution inputs. To remedy this drawback, this paper proposes a novel efficient attention mechanism equivalent to dot-product attention but with substantially less memory and computational costs. Its resource efficiency allows more widespread and flexible integration of attention modules into a network, which leads to better accuracies. Empirical evaluations demonstrated the effectiveness of its advantages. Efficient attention modules brought significant performance boosts to object detectors and instance segmenters on MS-COCO 2017. Further, the resource efficiency democratizes attention to complex models, where high costs prohibit the use of dot-product attention. As an exemplar, a model with efficient attention achieved state-of-the-art accuracies for stereo depth estimation on the Scene Flow dataset. Code is available at https://github.com/cmsflash/efficient-attention.","url_abs":"https://arxiv.org/abs/1812.01243v9","url_pdf":"https://arxiv.org/pdf/1812.01243v9.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/cmsflash/efficient-attention","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/96jonesa/StyleGan2-Colab-Demo","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/Bakikii/stylegan2-pytorch23","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/Di-Is/stylegan2-ada-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/HighCWu/stylegan2-paddle","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/ShreyasArthur/StyleGAN-2-with-Urban-Plans","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/SiavashCS/sgan_simple","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/lucidrains/DALLE2-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/lucidrains/En-transformer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/lucidrains/linear-attention-transformer","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/lucidrains/memory-transformer-xl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/lucidrains/stylegan2-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/MindCode-4/code-7/tree/main/factorized-attention","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"factorized-attention-self-attention-with","repo_url":"https://github.com/guluguluhhhh/contrib/tree/43-1/application/factorized-attention-self-attention-with-linear-complexities","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"extractive-document-summarization","task_name":"Extractive Text Summarization"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-recognition","task_name":"Object Recognition"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"stereo-depth-estimation","task_name":"Stereo Depth Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/extractive-text-summarization-on-govreport","task":"Extractive Text Summarization","dataset":"GovReport","model":"HEPOS","rank_in_archive_order":2,"of":2,"metrics":{"Avg. Test Rouge1":"56.86","Avg. Test Rouge2":"22.62","Avg. Test RougeLsum":"53.82"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.01243","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1812.01243"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/guluguluhhhh/contrib/tree/43-1/application/factorized-attention-self-attention-with-linear-complexities","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindCode-4/code-7/tree/main/factorized-attention","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/HighCWu/stylegan2-paddle","reach":{"status":"ok","spdx":"GPL-3.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lucidrains/stylegan2-pytorch","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Bakikii/stylegan2-pytorch23","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/96jonesa/StyleGan2-Colab-Demo","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/SiavashCS/sgan_simple","reach":{"status":"ok","spdx":"GPL-3.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lucidrains/linear-attention-transformer","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ShreyasArthur/StyleGAN-2-with-Urban-Plans","reach":{"status":"ok","spdx":"GPL-3.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/cmsflash/efficient-attention","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lucidrains/DALLE2-pytorch","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lucidrains/En-transformer","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Di-Is/stylegan2-ada-pytorch","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lucidrains/memory-transformer-xl","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_draft_wrong":3,"ran_violates":4,"unverified":1},"by_repo_kind":{"listed":{"samples":7,"ran":6,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"dd3db53ae7ad1c85","entry":"always","repo":"lucidrains/linear-attention-transformer","repo_kind":"listed","path":"linear_attention_transformer/linear_attention_transformer.py","file_url":"https://github.com/lucidrains/linear-attention-transformer/blob/HEAD/linear_attention_transformer/linear_attention_transformer.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"dd3db53ae7ad1c85"}},{"code_sha256_prefix":"369eced62e96d690","entry":"cast_tuple","repo":"lucidrains/memory-transformer-xl","repo_kind":"listed","path":"memory_transformer_xl/memory_transformer_xl.py","file_url":"https://github.com/lucidrains/memory-transformer-xl/blob/HEAD/memory_transformer_xl/memory_transformer_xl.py","link_basis":"harvester_set","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"369eced62e96d690"}},{"code_sha256_prefix":"a1bda7590dd9a4d2","entry":"default","repo":"lucidrains/memory-transformer-xl","repo_kind":"listed","path":"memory_transformer_xl/memory_transformer_xl.py","file_url":"https://github.com/lucidrains/memory-transformer-xl/blob/HEAD/memory_transformer_xl/memory_transformer_xl.py","link_basis":"harvester_set","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a1bda7590dd9a4d2"}},{"code_sha256_prefix":"b5307f70349f5c27","entry":"default","repo":"lucidrains/linear-attention-transformer","repo_kind":"listed","path":"linear_attention_transformer/linear_attention_transformer.py","file_url":"https://github.com/lucidrains/linear-attention-transformer/blob/HEAD/linear_attention_transformer/linear_attention_transformer.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b5307f70349f5c27"}},{"code_sha256_prefix":"aa5486a3650902d8","entry":"exists","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"aa5486a3650902d8"}},{"code_sha256_prefix":"7417c98099ec4e15","entry":"to","repo":"lucidrains/memory-transformer-xl","repo_kind":"listed","path":"memory_transformer_xl/memory_transformer_xl.py","file_url":"https://github.com/lucidrains/memory-transformer-xl/blob/HEAD/memory_transformer_xl/memory_transformer_xl.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7417c98099ec4e15"}},{"code_sha256_prefix":"3b65d5d77510c964","entry":"top_k","repo":"lucidrains/memory-transformer-xl","repo_kind":"listed","path":"memory_transformer_xl/autoregressive_wrapper.py","file_url":"https://github.com/lucidrains/memory-transformer-xl/blob/HEAD/memory_transformer_xl/autoregressive_wrapper.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3b65d5d77510c964"}},{"code_sha256_prefix":"864f3eecd1ea0a36","entry":"top_p","repo":"lucidrains/memory-transformer-xl","repo_kind":"listed","path":"memory_transformer_xl/autoregressive_wrapper.py","file_url":"https://github.com/lucidrains/memory-transformer-xl/blob/HEAD/memory_transformer_xl/autoregressive_wrapper.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"864f3eecd1ea0a36"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}