{"url":"/method/sparse-transformer","slug":"sparse-transformer","name":"Sparse Transformer","full_name":"Sparse Transformer","full_name_withheld":false,"description_markdown":"A **Sparse Transformer** is a [Transformer](https://paperswithcode.com/method/transformer) based architecture which utilises sparse factorizations of the attention matrix to reduce time/memory to $O(n \\sqrt{n})$. Other changes to the Transformer architecture include: (a) a restructured [residual block](https://paperswithcode.com/method/residual-block) and weight initialization, (b) A set of sparse attention kernels which efficiently compute subsets of the attention matrix, (c) recomputation of attention weights during the backwards pass to reduce memory usage","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"http://arxiv.org/abs/1904.10509v1","title":"Generating Long Sequences with Sparse Transformers","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":47,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"Pyramid Sparse Transformer: Enhancing Multi-Scale Feature Fusion with Dynamic Token Selection","date":"2025-05-19","arxiv_id":"2505.12772","n_code_links":0,"syntology":null},{"paper":null,"title":"Tractable Transformers for Flexible Conditional Generation","date":"2025-02-11","arxiv_id":"2502.07616","n_code_links":0,"syntology":null},{"paper":"/paper/deepgate4-efficient-and-effective","title":"DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale","date":"2025-02-02","arxiv_id":"2502.01681","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":1}},{"paper":null,"title":"SPARTAN: A Sparse Transformer Learning Local Causation","date":"2024-11-11","arxiv_id":"2411.06890","n_code_links":0,"syntology":null},{"paper":null,"title":"Extra Global Attention Designation Using Keyword Detection in Sparse Transformer Architectures","date":"2024-10-11","arxiv_id":"2410.08971","n_code_links":0,"syntology":null},{"paper":"/paper/efficient-and-scalable-point-cloud-generation","title":"Efficient and Scalable Point Cloud Generation with Sparse Point-Voxel Diffusion Models","date":"2024-08-12","arxiv_id":"2408.06145","n_code_links":2,"syntology":null},{"paper":"/paper/stretching-each-dollar-diffusion-training","title":"Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget","date":"2024-07-22","arxiv_id":"2407.15811","n_code_links":1,"syntology":{"ran":4,"of":6,"unverified":2,"pointer_only":0}},{"paper":null,"title":"HDT: Hierarchical Document Transformer","date":"2024-07-11","arxiv_id":"2407.08330","n_code_links":0,"syntology":null},{"paper":"/paper/fibottention-inceptive-visual-representation","title":"Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads","date":"2024-06-27","arxiv_id":"2406.19391","n_code_links":1,"syntology":null},{"paper":"/paper/a-mixture-of-experts-approach-to-3d-human","title":"A Mixture of Experts Approach to 3D Human Motion Prediction","date":"2024-05-09","arxiv_id":"2405.06088","n_code_links":1,"syntology":null},{"paper":"/paper/scene-adaptive-sparse-transformer-for-event","title":"Scene Adaptive Sparse Transformer for Event-based Object Detection","date":"2024-04-02","arxiv_id":"2404.01882","n_code_links":1,"syntology":{"ran":19,"of":23,"unverified":4,"pointer_only":0}},{"paper":"/paper/smtf-sparse-transformer-with-multiscale","title":"SMTF: Sparse transformer with multiscale contextual fusion for medical image segmentation","date":"2024-03-24","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Segmentation Guided Sparse Transformer for Under-Display Camera Image Restoration","date":"2024-03-09","arxiv_id":"2403.05906","n_code_links":0,"syntology":null},{"paper":null,"title":"Do Efficient Transformers Really Save Computation?","date":"2024-02-21","arxiv_id":"2402.13934","n_code_links":0,"syntology":null},{"paper":"/paper/adapt-or-perish-adaptive-sparse-transformer","title":"Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image Restoration","date":"2024-01-01","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Towards More Unified In-context Visual Understanding","date":"2023-12-05","arxiv_id":"2312.02520","n_code_links":0,"syntology":null},{"paper":null,"title":"SPION: Layer-Wise Sparse Training of Transformer via Convolutional Flood Filling","date":"2023-09-22","arxiv_id":"2309.12578","n_code_links":0,"syntology":null},{"paper":"/paper/sparseswin-swin-transformer-with-sparse","title":"SparseSwin: Swin Transformer with Sparse Transformer Block","date":"2023-09-11","arxiv_id":"2309.05224","n_code_links":1,"syntology":null},{"paper":"/paper/from-sparse-to-soft-mixtures-of-experts","title":"From Sparse to Soft Mixtures of Experts","date":"2023-08-02","arxiv_id":"2308.00951","n_code_links":5,"syntology":{"ran":29,"of":33,"unverified":4,"pointer_only":8}},{"paper":"/paper/longcoder-a-long-range-pre-trained-language","title":"LongCoder: A Long-Range Pre-trained Language Model for Code Completion","date":"2023-06-26","arxiv_id":"2306.14893","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}},{"paper":"/paper/learning-a-sparse-transformer-network-for","title":"Learning A Sparse Transformer Network for Effective Image Deraining","date":"2023-03-21","arxiv_id":"2303.11950","n_code_links":1,"syntology":{"ran":8,"of":11,"unverified":3,"pointer_only":11}},{"paper":null,"title":"Sampled Transformer for Point Sets","date":"2023-02-28","arxiv_id":"2302.14346","n_code_links":0,"syntology":null},{"paper":"/paper/generating-a-structured-summary-of-numerous","title":"Generating a Structured Summary of Numerous Academic Papers: Dataset and Method","date":"2023-02-09","arxiv_id":"2302.04580","n_code_links":1,"syntology":null},{"paper":"/paper/diffuser-efficient-transformers-with-multi","title":"Diffuser: Efficient Transformers with Multi-hop Attention Diffusion for Long Sequences","date":"2022-10-21","arxiv_id":"2210.11794","n_code_links":1,"syntology":null},{"paper":"/paper/efficient-quantized-sparse-matrix-operations","title":"Efficient Quantized Sparse Matrix Operations on Tensor Cores","date":"2022-09-14","arxiv_id":"2209.06979","n_code_links":1,"syntology":null},{"paper":"/paper/dynast-dynamic-sparse-transformer-for","title":"DynaST: Dynamic Sparse Transformer for Exemplar-Guided Image Generation","date":"2022-07-13","arxiv_id":"2207.06124","n_code_links":1,"syntology":{"ran":5,"of":8,"unverified":3,"pointer_only":0}},{"paper":"/paper/sparsetir-composable-abstractions-for-sparse","title":"SparseTIR: Composable Abstractions for Sparse Compilation in Deep Learning","date":"2022-07-11","arxiv_id":"2207.04606","n_code_links":2,"syntology":{"ran":4,"of":6,"unverified":2,"pointer_only":0}},{"paper":"/paper/deepspeed-inference-enabling-efficient","title":"DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale","date":"2022-06-30","arxiv_id":"2207.00032","n_code_links":2,"syntology":null},{"paper":"/paper/video-sparse-transformer-with-attention","title":"Video Sparse Transformer With Attention-Guided Memory for Video Object Detection","date":"2022-06-17","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/what-dense-graph-do-you-need-for-self","title":"What Dense Graph Do You Need for Self-Attention?","date":"2022-05-27","arxiv_id":"2205.14014","n_code_links":1,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":6},{"task":"/task/decoder","name":"Decoder","papers":5},{"task":"/task/language-modeling","name":"Language Modeling","papers":4},{"task":"/task/mixture-of-experts","name":"Mixture-of-Experts","papers":4},{"task":"/task/text-classification","name":"Text Classification","papers":4},{"task":"/task/text-classification-1","name":"text-classification","papers":4},{"task":"/task/diversity","name":"Diversity","papers":3},{"task":null,"name":"GPU","papers":3},{"task":"/task/image-restoration","name":"Image Restoration","papers":3},{"task":"/task/machine-translation","name":"Machine Translation","papers":3},{"task":"/task/object","name":"Object","papers":3},{"task":"/task/object-detection","name":"Object Detection","papers":3},{"task":"/task/question-answering","name":"Question Answering","papers":3},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":3},{"task":"/task/translation","name":"Translation","papers":3},{"task":"/task/object-detection-1","name":"object-detection","papers":3},{"task":null,"name":"CPU","papers":2},{"task":"/task/clustering","name":"Clustering","papers":2},{"task":"/task/document-summarization","name":"Document Summarization","papers":2},{"task":"/task/image-captioning","name":"Image Captioning","papers":2}],"tasks_shown":20,"n_tasks":80,"usage_by_year":[{"year":"2019","papers":3},{"year":"2020","papers":5},{"year":"2021","papers":6},{"year":"2022","papers":10},{"year":"2023","papers":8},{"year":"2024","papers":12},{"year":"2025","papers":3}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/sparse-transformer"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}