{"url":"/method/transformer-xl","slug":"transformer-xl","name":"Transformer-XL","full_name":"Transformer-XL","full_name_withheld":false,"description_markdown":"**Transformer-XL** (meaning extra long) is a [Transformer](https://paperswithcode.com/method/transformer) architecture that introduces the notion of recurrence to the deep self-attention network. Instead of computing the hidden states from scratch for each new segment, Transformer-XL reuses the hidden states obtained in previous segments. The reused hidden states serve as memory for the current segment, which builds up a recurrent connection between the segments. As a result, modeling very long-term dependency becomes possible because information can be propagated through the recurrent connections. As an additional contribution, the Transformer-XL uses a new relative positional encoding formulation that generalizes to attention lengths longer than the one observed during training.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context","paper":"/paper/transformer-xl-attentive-language-models","first_author":"Zihang Dai","n_authors":6,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/transformer-xl-attentive-language-models"},"source":{"url":"https://arxiv.org/abs/1901.02860v3","title":"Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Autoregressive Transformers","url":"/methods/category/autoregressive-transformers","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":64,"archive_num_papers":64,"papers_newest_first":[{"paper":"/paper/rlbenchnet-the-right-network-for-the-right","title":"RLBenchNet: The Right Network for the Right Reinforcement Learning Task","date":"2025-05-21","arxiv_id":"2505.15040","n_code_links":1,"syntology":null},{"paper":null,"title":"A Combined Encoder and Transformer Approach for Coherent and High-Quality Text Generation","date":"2024-11-19","arxiv_id":"2411.12157","n_code_links":0,"syntology":null},{"paper":null,"title":"Large Body Language Models","date":"2024-10-21","arxiv_id":"2410.16533","n_code_links":0,"syntology":null},{"paper":null,"title":"Transformers for Supervised Online Continual Learning","date":"2024-03-03","arxiv_id":"2403.01554","n_code_links":0,"syntology":null},{"paper":"/paper/unimem-towards-a-unified-view-of-long-context","title":"UniMem: Towards a Unified View of Long-Context Large Language Models","date":"2024-02-05","arxiv_id":"2402.03009","n_code_links":1,"syntology":null},{"paper":"/paper/memory-efficient-stochastic-methods-for","title":"Memory-efficient Stochastic methods for Memory-based Transformers","date":"2023-11-14","arxiv_id":"2311.08123","n_code_links":1,"syntology":null},{"paper":"/paper/trams-training-free-memory-selection-for-long","title":"TRAMS: Training-free Memory Selection for Long-range Language Modeling","date":"2023-10-24","arxiv_id":"2310.15494","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}},{"paper":"/paper/approximating-two-layer-feedforward-networks","title":"Approximating Two-Layer Feedforward Networks for Efficient Transformers","date":"2023-10-16","arxiv_id":"2310.10837","n_code_links":2,"syntology":{"ran":3,"of":4,"unverified":1,"pointer_only":0}},{"paper":"/paper/memory-gym-partially-observable-challenges-to","title":"Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of Agents","date":"2023-09-29","arxiv_id":"2309.17207","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}},{"paper":"/paper/random-access-infinite-context-length-for","title":"Random-Access Infinite Context Length for Transformers","date":"2023-09-21","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/rcmha-relative-convolutional-multi-head","title":"RCMHA: Relative Convolutional Multi-Head Attention for Natural Language Modelling","date":"2023-08-07","arxiv_id":"2308.03429","n_code_links":1,"syntology":null},{"paper":"/paper/landmark-attention-random-access-infinite","title":"Landmark Attention: Random-Access Infinite Context Length for Transformers","date":"2023-05-25","arxiv_id":"2305.16300","n_code_links":2,"syntology":{"ran":11,"of":13,"unverified":2,"pointer_only":0}},{"paper":"/paper/transformer-based-world-models-are-happy-with","title":"Transformer-based World Models Are Happy With 100k Interactions","date":"2023-03-13","arxiv_id":"2303.07109","n_code_links":1,"syntology":{"ran":16,"of":25,"unverified":9,"pointer_only":0}},{"paper":null,"title":"GTR-CTRL: Instrument and Genre Conditioning for Guitar-Focused Music Generation with Transformers","date":"2023-02-10","arxiv_id":"2302.05393","n_code_links":0,"syntology":null},{"paper":null,"title":"An Comparative Analysis of Different Pitch and Metrical Grid Encoding Methods in the Task of Sequential Music Generation","date":"2023-01-31","arxiv_id":"2301.13383","n_code_links":0,"syntology":null},{"paper":null,"title":"Efficient Sparsely Activated Transformers","date":"2022-08-31","arxiv_id":"2208.14580","n_code_links":0,"syntology":null},{"paper":"/paper/adan-adaptive-nesterov-momentum-algorithm-for","title":"Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models","date":"2022-08-13","arxiv_id":"2208.06677","n_code_links":9,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}},{"paper":"/paper/recurrent-memory-transformer","title":"Recurrent Memory Transformer","date":"2022-07-14","arxiv_id":"2207.06881","n_code_links":3,"syntology":{"ran":6,"of":13,"unverified":7,"pointer_only":2}},{"paper":"/paper/emotion-aware-transformer-encoder-for-1","title":"Emotion-Aware Transformer Encoder for Empathetic Dialogue Generation","date":"2022-04-24","arxiv_id":"2204.11320","n_code_links":1,"syntology":null},{"paper":"/paper/sintra-learning-an-inspiration-model-from-a","title":"SinTra: Learning an inspiration model from a single multi-track music segment","date":"2022-04-21","arxiv_id":"2204.09917","n_code_links":1,"syntology":null},{"paper":"/paper/litetransformersearch-training-free-on-device","title":"LiteTransformerSearch: Training-free Neural Architecture Search for Efficient Language Models","date":"2022-03-04","arxiv_id":"2203.02094","n_code_links":1,"syntology":{"ran":1,"of":5,"unverified":4,"pointer_only":0}},{"paper":null,"title":"Reconsidering the Past: Optimizing Hidden States in Language Models","date":"2021-12-16","arxiv_id":"2112.08653","n_code_links":0,"syntology":null},{"paper":null,"title":"A Comparative Study of Transformers on Word Sense Disambiguation","date":"2021-11-30","arxiv_id":"2111.15417","n_code_links":0,"syntology":null},{"paper":null,"title":"How much do language models copy from their training data? Evaluating linguistic novelty in text generation using RAVEN","date":"2021-11-18","arxiv_id":"2111.09509","n_code_links":0,"syntology":null},{"paper":"/paper/synthesizing-collective-communication","title":"TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches","date":"2021-11-08","arxiv_id":"2111.04867","n_code_links":2,"syntology":null},{"paper":null,"title":"Language Modelling via Learning to Rank","date":"2021-10-13","arxiv_id":"2110.06961","n_code_links":0,"syntology":null},{"paper":"/paper/layer-wise-pruning-of-transformer-attention","title":"Layer-wise Pruning of Transformer Attention Heads for Efficient Language Modeling","date":"2021-10-07","arxiv_id":"2110.03252","n_code_links":1,"syntology":null},{"paper":"/paper/fnetar-mixing-tokens-with-autoregressive","title":"FNetAR: Mixing Tokens with Autoregressive Fourier Transforms","date":"2021-07-22","arxiv_id":"2107.10932","n_code_links":1,"syntology":null},{"paper":"/paper/transformers-with-multi-modal-features-and","title":"Transformers with multi-modal features and post-fusion context for e-commerce session-based recommendation","date":"2021-07-11","arxiv_id":"2107.05124","n_code_links":0,"syntology":null},{"paper":null,"title":"ASR Adaptation for E-commerce Chatbots using Cross-Utterance Context and Multi-Task Language Modeling","date":"2021-06-15","arxiv_id":"2106.09532","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":38},{"task":"/task/language-modeling","name":"Language Modeling","papers":27},{"task":"/task/decoder","name":"Decoder","papers":7},{"task":"/task/machine-translation","name":"Machine Translation","papers":6},{"task":"/task/speech-recognition","name":"Speech Recognition","papers":5},{"task":"/task/translation","name":"Translation","papers":5},{"task":"/task/speech-recognition-1","name":"speech-recognition","papers":5},{"task":"/task/text-generation","name":"Text Generation","papers":4},{"task":"/task/automatic-speech-recognition-2","name":"Automatic Speech Recognition","papers":3},{"task":"/task/automatic-speech-recognition","name":"Automatic Speech Recognition (ASR)","papers":3},{"task":"/task/paraphrase-identification","name":"Paraphrase Identification","papers":3},{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":3},{"task":"/task/sentence","name":"Sentence","papers":3},{"task":"/task/word-embeddings","name":"Word Embeddings","papers":3},{"task":"/task/abstractive-text-summarization","name":"Abstractive Text Summarization","papers":2},{"task":"/task/deep-attention","name":"Deep Attention","papers":2},{"task":"/task/deep-reinforcement-learning","name":"Deep Reinforcement Learning","papers":2},{"task":null,"name":"GPU","papers":2},{"task":"/task/graph-neural-network","name":"Graph Neural Network","papers":2},{"task":"/task/music-generation","name":"Music Generation","papers":2}],"tasks_shown":20,"n_tasks":86,"usage_by_year":[{"year":"2019","papers":13},{"year":"2020","papers":16},{"year":"2021","papers":14},{"year":"2022","papers":6},{"year":"2023","papers":10},{"year":"2024","papers":4},{"year":"2025","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/transformer-xl"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}