Methods › Natural Language Processing › Autoregressive Transformers › Transformer-XL
Transformer-XL
Introduced by Zihang Dai et al. in Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Transformer-XL (meaning extra long) is a Transformer architecture that introduces the notion of recurrence to the deep self-attention network. Instead of computing the hidden states from scratch for each new segment, Transformer-XL reuses the hidden states obtained in previous segments. The reused hidden states serve as memory for the current segment, which builds up a recurrent connection between the segments. As a result, modeling very long-term dependency becomes possible because information can be propagated through the recurrent connections. As an additional contribution, the Transformer-XL uses a new relative positional encoding formulation that generalizes to attention lengths longer than the one observed during training.
Papers archive 2025-07-28
30 shown of 64, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
RLBenchNet: The Right Network for the Right Reinforcement Learning Task 21 May 2025 · 1 repository · arXiv:2505.15040
-
A Combined Encoder and Transformer Approach for Coherent and High-Quality Text Generation 19 Nov 2024 · 0 repositories · arXiv:2411.12157
-
Large Body Language Models 21 Oct 2024 · 0 repositories · arXiv:2410.16533
-
Transformers for Supervised Online Continual Learning 3 Mar 2024 · 0 repositories · arXiv:2403.01554
-
UniMem: Towards a Unified View of Long-Context Large Language Models 5 Feb 2024 · 1 repository · arXiv:2402.03009
-
Memory-efficient Stochastic methods for Memory-based Transformers 14 Nov 2023 · 1 repository · arXiv:2311.08123
-
TRAMS: Training-free Memory Selection for Long-range Language Modeling 24 Oct 2023 · 1 repository · arXiv:2310.15494Syntology ran 1 of 1 samples · 0 unverified
-
Approximating Two-Layer Feedforward Networks for Efficient Transformers 16 Oct 2023 · 2 repositories · arXiv:2310.10837Syntology ran 3 of 4 samples · 1 unverified
-
Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of Agents 29 Sep 2023 · 1 repository · arXiv:2309.17207Syntology ran 1 of 1 samples · 0 unverified
-
Random-Access Infinite Context Length for Transformers 21 Sep 2023 · 1 repository
-
RCMHA: Relative Convolutional Multi-Head Attention for Natural Language Modelling 7 Aug 2023 · 1 repository · arXiv:2308.03429
-
Landmark Attention: Random-Access Infinite Context Length for Transformers 25 May 2023 · 2 repositories · arXiv:2305.16300Syntology ran 11 of 13 samples · 2 unverified
-
Transformer-based World Models Are Happy With 100k Interactions 13 Mar 2023 · 1 repository · arXiv:2303.07109Syntology ran 16 of 25 samples · 9 unverified
-
GTR-CTRL: Instrument and Genre Conditioning for Guitar-Focused Music Generation with Transformers 10 Feb 2023 · 0 repositories · arXiv:2302.05393
-
An Comparative Analysis of Different Pitch and Metrical Grid Encoding Methods in the Task of Sequential Music Generation 31 Jan 2023 · 0 repositories · arXiv:2301.13383
-
Efficient Sparsely Activated Transformers 31 Aug 2022 · 0 repositories · arXiv:2208.14580
-
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models 13 Aug 2022 · 9 repositories · arXiv:2208.06677Syntology ran 1 of 1 samples · 0 unverified
-
Recurrent Memory Transformer 14 Jul 2022 · 3 repositories · arXiv:2207.06881Syntology ran 6 of 13 samples · 7 unverified · 2 pointer-only (licence)
-
Emotion-Aware Transformer Encoder for Empathetic Dialogue Generation 24 Apr 2022 · 1 repository · arXiv:2204.11320
-
SinTra: Learning an inspiration model from a single multi-track music segment 21 Apr 2022 · 1 repository · arXiv:2204.09917
-
LiteTransformerSearch: Training-free Neural Architecture Search for Efficient Language Models 4 Mar 2022 · 1 repository · arXiv:2203.02094Syntology ran 1 of 5 samples · 4 unverified
-
Reconsidering the Past: Optimizing Hidden States in Language Models 16 Dec 2021 · 0 repositories · arXiv:2112.08653
-
A Comparative Study of Transformers on Word Sense Disambiguation 30 Nov 2021 · 0 repositories · arXiv:2111.15417
-
How much do language models copy from their training data? Evaluating linguistic novelty in text generation using RAVEN 18 Nov 2021 · 0 repositories · arXiv:2111.09509
-
TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches 8 Nov 2021 · 2 repositories · arXiv:2111.04867
-
Language Modelling via Learning to Rank 13 Oct 2021 · 0 repositories · arXiv:2110.06961
-
Layer-wise Pruning of Transformer Attention Heads for Efficient Language Modeling 7 Oct 2021 · 1 repository · arXiv:2110.03252
-
FNetAR: Mixing Tokens with Autoregressive Fourier Transforms 22 Jul 2021 · 1 repository · arXiv:2107.10932
-
Transformers with multi-modal features and post-fusion context for e-commerce session-based recommendation 11 Jul 2021 · 0 repositories · arXiv:2107.05124
-
ASR Adaptation for E-commerce Chatbots using Cross-Utterance Context and Multi-Task Language Modeling 15 Jun 2021 · 0 repositories · arXiv:2106.09532
Tasks archive 2025-07-28
20 shown of 86 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections