Methods › Natural Language Processing › Attention Patterns › BigBird
BigBird
Introduced by Manzil Zaheer et al. in Big Bird: Transformers for Longer Sequences
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
BigBird is a Transformer with a sparse attention mechanism that reduces the quadratic dependency of self-attention to linear in the number of tokens. BigBird is a universal approximator of sequence functions and is Turing complete, thereby preserving these properties of the quadratic, full attention model. In particular, BigBird consists of three main parts:
- A set of g global tokens attending on all parts of the sequence.
- All tokens attending to a set of w local neighboring tokens.
- All tokens attending to a set of r random tokens.
This leads to a high performing attention mechanism scaling to much longer sequence lengths (8x).
Papers archive 2025-07-28
16 shown of 16, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Convolutional vs Large Language Models for Software Log Classification in Edge-Deployable Cellular Network Testing 4 Jul 2024 · 0 repositories · arXiv:2407.03759
-
Transfer Learning in Pre-Trained Large Language Models for Malware Detection Based on System Calls 15 May 2024 · 0 repositories · arXiv:2405.09318
-
Multi-level Contrastive Learning for Script-based Character Understanding 20 Oct 2023 · 1 repository · arXiv:2310.13231
-
KoBigBird-large: Transformation of Transformer for Korean Language Understanding 19 Sep 2023 · 0 repositories · arXiv:2309.10339
-
BudgetLongformer: Can we Cheaply Pretrain a SotA Legal Language Model From Scratch? 30 Nov 2022 · 0 repositories · arXiv:2211.17135
-
Processing Long Legal Documents with Pre-trained Transformers: Modding LegalBERT and Longformer 2 Nov 2022 · 0 repositories · arXiv:2211.00974
-
LittleBird: Efficient Faster & Longer Transformer for Question Answering 21 Oct 2022 · 0 repositories · arXiv:2210.11870
-
Factorizing Content and Budget Decisions in Abstractive Summarization of Long Documents 25 May 2022 · 1 repository · arXiv:2205.12486
-
ICDBigBird: A Contextual Embedding Model for ICD Code Classification 21 Apr 2022 · 0 repositories · arXiv:2204.10408
-
Clinical-Longformer and Clinical-BigBird: Transformers for long clinical sequences 27 Jan 2022 · 1 repository · arXiv:2201.11838
-
Hierarchical Neural Network Approaches for Long Document Classification 18 Jan 2022 · 0 repositories · arXiv:2201.06774
-
Dynamic Token Normalization Improves Vision Transformers 5 Dec 2021 · 1 repository · arXiv:2112.02624Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)
-
Hierarchical Transformer Networks for Long-sequence and Multiple Clinical Documents Classification 16 Oct 2021 · 0 repositories
-
A Dataset for Answering Time-Sensitive Questions 13 Aug 2021 · 1 repository · arXiv:2108.06314Syntology ran 4 of 4 samples · 0 unverified
-
Three-level Hierarchical Transformer Networks for Long-sequence and Multiple Clinical Documents Classification 17 Apr 2021 · 1 repository · arXiv:2104.08444
-
Big Bird: Transformers for Longer Sequences 28 Jul 2020 · 14 repositories · arXiv:2007.14062Syntology ran 10 of 15 samples · 5 unverified · 11 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 35 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections