Methods › Natural Language Processing › Transformers › Sparse Transformer
Sparse Transformer
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
A Sparse Transformer is a Transformer based architecture which utilises sparse factorizations of the attention matrix to reduce time/memory to O(n √(n)). Other changes to the Transformer architecture include: (a) a restructured residual block and weight initialization, (b) A set of sparse attention kernels which efficiently compute subsets of the attention matrix, (c) recomputation of attention weights during the backwards pass to reduce memory usage
Papers archive 2025-07-28
30 shown of 47, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Pyramid Sparse Transformer: Enhancing Multi-Scale Feature Fusion with Dynamic Token Selection 19 May 2025 · 0 repositories · arXiv:2505.12772
-
Tractable Transformers for Flexible Conditional Generation 11 Feb 2025 · 0 repositories · arXiv:2502.07616
-
DeepGate4: Efficient and Effective Representation Learning for Circuit Design at Scale 2 Feb 2025 · 1 repository · arXiv:2502.01681Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)
-
SPARTAN: A Sparse Transformer Learning Local Causation 11 Nov 2024 · 0 repositories · arXiv:2411.06890
-
Extra Global Attention Designation Using Keyword Detection in Sparse Transformer Architectures 11 Oct 2024 · 0 repositories · arXiv:2410.08971
-
Efficient and Scalable Point Cloud Generation with Sparse Point-Voxel Diffusion Models 12 Aug 2024 · 2 repositories · arXiv:2408.06145
-
Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget 22 Jul 2024 · 1 repository · arXiv:2407.15811Syntology ran 4 of 6 samples · 2 unverified
-
HDT: Hierarchical Document Transformer 11 Jul 2024 · 0 repositories · arXiv:2407.08330
-
Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads 27 Jun 2024 · 1 repository · arXiv:2406.19391
-
A Mixture of Experts Approach to 3D Human Motion Prediction 9 May 2024 · 1 repository · arXiv:2405.06088
-
Scene Adaptive Sparse Transformer for Event-based Object Detection 2 Apr 2024 · 1 repository · arXiv:2404.01882Syntology ran 19 of 23 samples · 4 unverified
-
SMTF: Sparse transformer with multiscale contextual fusion for medical image segmentation 24 Mar 2024 · 1 repository
-
Segmentation Guided Sparse Transformer for Under-Display Camera Image Restoration 9 Mar 2024 · 0 repositories · arXiv:2403.05906
-
Do Efficient Transformers Really Save Computation? 21 Feb 2024 · 0 repositories · arXiv:2402.13934
-
Adapt or Perish: Adaptive Sparse Transformer with Attentive Feature Refinement for Image Restoration 1 Jan 2024 · 1 repository
-
Towards More Unified In-context Visual Understanding 5 Dec 2023 · 0 repositories · arXiv:2312.02520
-
SPION: Layer-Wise Sparse Training of Transformer via Convolutional Flood Filling 22 Sep 2023 · 0 repositories · arXiv:2309.12578
-
SparseSwin: Swin Transformer with Sparse Transformer Block 11 Sep 2023 · 1 repository · arXiv:2309.05224
-
From Sparse to Soft Mixtures of Experts 2 Aug 2023 · 5 repositories · arXiv:2308.00951Syntology ran 29 of 33 samples · 4 unverified · 8 pointer-only (licence)
-
LongCoder: A Long-Range Pre-trained Language Model for Code Completion 26 Jun 2023 · 1 repository · arXiv:2306.14893Syntology ran 1 of 1 samples · 0 unverified
-
Learning A Sparse Transformer Network for Effective Image Deraining 21 Mar 2023 · 1 repository · arXiv:2303.11950Syntology ran 8 of 11 samples · 3 unverified · 11 pointer-only (licence)
-
Sampled Transformer for Point Sets 28 Feb 2023 · 0 repositories · arXiv:2302.14346
-
Generating a Structured Summary of Numerous Academic Papers: Dataset and Method 9 Feb 2023 · 1 repository · arXiv:2302.04580
-
Diffuser: Efficient Transformers with Multi-hop Attention Diffusion for Long Sequences 21 Oct 2022 · 1 repository · arXiv:2210.11794
-
Efficient Quantized Sparse Matrix Operations on Tensor Cores 14 Sep 2022 · 1 repository · arXiv:2209.06979
-
DynaST: Dynamic Sparse Transformer for Exemplar-Guided Image Generation 13 Jul 2022 · 1 repository · arXiv:2207.06124Syntology ran 5 of 8 samples · 3 unverified
-
SparseTIR: Composable Abstractions for Sparse Compilation in Deep Learning 11 Jul 2022 · 2 repositories · arXiv:2207.04606Syntology ran 4 of 6 samples · 2 unverified
-
DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale 30 Jun 2022 · 2 repositories · arXiv:2207.00032
-
Video Sparse Transformer With Attention-Guided Memory for Video Object Detection 17 Jun 2022 · 1 repository
-
What Dense Graph Do You Need for Self-Attention? 27 May 2022 · 1 repository · arXiv:2205.14014
Tasks archive 2025-07-28
20 shown of 80 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections