Methods › Natural Language Processing › Transformers › Compressive Transformer
Compressive Transformer
Introduced by Jack W. Rae et al. in Compressive Transformers for Long-Range Sequence Modelling
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
The Compressive Transformer is an extension to the Transformer which maps past hidden activations (memories) to a smaller set of compressed representations (compressed memories). The Compressive Transformer uses the same attention mechanism over its set of memories and compressed memories, learning to query both its short-term granular memory and longer-term coarse memory. It builds on the ideas of Transformer-XL which maintains a memory of past activations at each layer to preserve a longer history of context. The Transformer-XL discards past activations when they become sufficiently old (controlled by the size of the memory). The key principle of the Compressive Transformer is to compress these old memories, instead of discarding them, and store them in an additional compressed memory.
At each time step t, we discard the oldest compressed memories (FIFO) and then the oldest n states from ordinary memory are compressed and shifted to the new slot in compressed memory. During training, the compressive memory component is optimized separately from the main language model (separate training loop).
Papers archive 2025-07-28
3 shown of 3, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Collection Space Navigator: An Interactive Visualization Interface for Multidimensional Datasets 11 May 2023 · 2 repositories · arXiv:2305.06809
-
DCT: Dynamic Compressive Transformer for Modeling Unbounded Sequence 10 Oct 2021 · 0 repositories · arXiv:2110.04821
-
Compressive Transformers for Long-Range Sequence Modelling 13 Nov 2019 · 6 repositories · arXiv:1911.05507Syntology ran 3 of 11 samples · 8 unverified
Tasks archive 2025-07-28
5 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| Data Visualization | 1 |
| Dimensionality Reduction | 1 |
| Embeddings Evaluation | 1 |
| Language Modelling | 1 |
| Sentence | 1 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections