Methods › Natural Language Processing › Autoencoding Transformers › Longformer
Longformer
Introduced by Iz Beltagy et al. in Longformer: The Long-Document Transformer
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Longformer is a modified Transformer architecture. Traditional Transformer-based models are unable to process long sequences due to their self-attention operation, which scales quadratically with the sequence length. To address this, Longformer uses an attention pattern that scales linearly with sequence length, making it easy to process documents of thousands of tokens or longer. The attention mechanism is a drop-in replacement for the standard self-attention and combines a local windowed attention with a task motivated global attention.
The attention patterns utilised include: sliding window attention, dilated sliding window attention and global + sliding window. These can be viewed in the components section of this page.
Papers archive 2025-07-28
30 shown of 87, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution 18 Jun 2025 · 0 repositories · arXiv:2506.17323
-
Enhancing Abstractive Summarization of Scientific Papers Using Structure Information 20 May 2025 · 1 repository · arXiv:2505.14179
-
CacheFormer: High Attention-Based Segment Caching 18 Apr 2025 · 0 repositories · arXiv:2504.13981
-
ARLED: Leveraging LED-based ARMAN Model for Abstractive Summarization of Persian Long Documents 13 Mar 2025 · 0 repositories · arXiv:2503.10233
-
Understanding Players as if They Are Talking to the Game in a Customized Language: A Pilot Study 24 Oct 2024 · 0 repositories · arXiv:2410.18605
-
Extra Global Attention Designation Using Keyword Detection in Sparse Transformer Architectures 11 Oct 2024 · 0 repositories · arXiv:2410.08971
-
The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models 9 Oct 2024 · 1 repository · arXiv:2410.06554Syntology ran 5 of 9 samples · 4 unverified
-
RICo: Reddit ideological communities 5 Jun 2024 · 1 repository
-
Transfer Learning in Pre-Trained Large Language Models for Malware Detection Based on System Calls 15 May 2024 · 0 repositories · arXiv:2405.09318
-
Advancing AI with Integrity: Ethical Challenges and Solutions in Neural Machine Translation 1 Apr 2024 · 0 repositories · arXiv:2404.01070
-
A multi-cohort study on prediction of acute brain dysfunction states using selective state space models 11 Mar 2024 · 0 repositories · arXiv:2403.07201
-
Adaptation of Biomedical and Clinical Pretrained Models to French Long Documents: A Comparative Study 26 Feb 2024 · 1 repository · arXiv:2402.16689
-
Accurate and Well-Calibrated ICD Code Assignment Through Attention Over Diverse Label Embeddings 5 Feb 2024 · 1 repository · arXiv:2402.03172Syntology ran 7 of 7 samples · 0 unverified
-
UniMem: Towards a Unified View of Long-Context Large Language Models 5 Feb 2024 · 1 repository · arXiv:2402.03009
-
Enhanced Labeling Technique for Reddit Text and Fine-Tuned Longformer Models for Classifying Depression Severity in English and Luganda 25 Jan 2024 · 0 repositories · arXiv:2401.14240
-
Exploring Automatic Text Simplification of German Narrative Documents 15 Dec 2023 · 1 repository · arXiv:2312.09907
-
LLVMs4Protest: Harnessing the Power of Large Language and Vision Models for Deciphering Protests in the News 30 Nov 2023 · 1 repository · arXiv:2311.18241
-
Towards Harmful Erotic Content Detection through Coreference-Driven Contextual Analysis 22 Oct 2023 · 0 repositories · arXiv:2310.14325
-
Multi-level Contrastive Learning for Script-based Character Understanding 20 Oct 2023 · 1 repository · arXiv:2310.13231
-
Improving Long Document Topic Segmentation Models With Enhanced Coherence Modeling 18 Oct 2023 · 1 repository · arXiv:2310.11772
-
Hallucination Reduction in Long Input Text Summarization 28 Sep 2023 · 1 repository · arXiv:2309.16781
-
An NLP Benchmark Dataset for Assessing Corporate Climate Policy Engagement 26 Sep 2023 · 0 repositories
-
Language Models for Novelty Detection in System Call Traces 5 Sep 2023 · 0 repositories · arXiv:2309.02206
-
Local Large Language Models for Complex Structured Medical Tasks 3 Aug 2023 · 1 repository · arXiv:2308.01727
-
Can Model Fusing Help Transformers in Long Document Classification? An Empirical Study 18 Jul 2023 · 1 repository · arXiv:2307.09532
-
NOWJ at COLIEE 2023 -- Multi-Task and Ensemble Approaches in Legal Information Processing 8 Jun 2023 · 0 repositories · arXiv:2306.04903
-
MultiLegalPile: A 689GB Multilingual Legal Corpus 3 Jun 2023 · 0 repositories · arXiv:2306.02069
-
DEPLAIN: A German Parallel Corpus with Intralingual Translations into Plain Language for Sentence and Document Simplification 30 May 2023 · 1 repository · arXiv:2305.18939
-
Incorporating Distributions of Discourse Structure for Long Document Abstractive Summarization 26 May 2023 · 0 repositories · arXiv:2305.16784
-
Neural Summarization of Electronic Health Records 24 May 2023 · 0 repositories · arXiv:2305.15222
Tasks archive 2025-07-28
20 shown of 111 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections