Methods › General › Output Functions › Adaptive Softmax
Adaptive Softmax
Introduced by Edouard Grave et al. in Efficient softmax approximation for GPUs
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Adaptive Softmax is a speedup technique for the computation of probability distributions over words. The adaptive softmax is inspired by the class-based hierarchical softmax, where the word classes are built to minimize the computation time. Adaptive softmax achieves efficiency by explicitly taking into account the computation time of matrix-multiplication on parallel systems and combining it with a few important observations, namely keeping a shortlist of frequent words in the root node and reducing the capacity of rare words.
Papers archive 2025-07-28
30 shown of 72, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
RLBenchNet: The Right Network for the Right Reinforcement Learning Task 21 May 2025 · 1 repository · arXiv:2505.15040
-
VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits 15 May 2025 · 0 repositories · arXiv:2505.10202
-
Convergence Rates for Softmax Gating Mixture of Experts 5 Mar 2025 · 0 repositories · arXiv:2503.03213
-
A Combined Encoder and Transformer Approach for Coherent and High-Quality Text Generation 19 Nov 2024 · 0 repositories · arXiv:2411.12157
-
Large Body Language Models 21 Oct 2024 · 0 repositories · arXiv:2410.16533
-
DenoMamba: A fused state-space model for low-dose CT denoising 19 Sep 2024 · 1 repository · arXiv:2409.13094
-
Online Residual Learning from Offline Experts for Pedestrian Tracking 6 Sep 2024 · 0 repositories · arXiv:2409.04069
-
Transformers for Supervised Online Continual Learning 3 Mar 2024 · 0 repositories · arXiv:2403.01554
-
UniMem: Towards a Unified View of Long-Context Large Language Models 5 Feb 2024 · 1 repository · arXiv:2402.03009
-
Memory-efficient Stochastic methods for Memory-based Transformers 14 Nov 2023 · 1 repository · arXiv:2311.08123
-
TRAMS: Training-free Memory Selection for Long-range Language Modeling 24 Oct 2023 · 1 repository · arXiv:2310.15494Syntology ran 1 of 1 samples · 0 unverified
-
Approximating Two-Layer Feedforward Networks for Efficient Transformers 16 Oct 2023 · 2 repositories · arXiv:2310.10837Syntology ran 3 of 4 samples · 1 unverified
-
Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of Agents 29 Sep 2023 · 1 repository · arXiv:2309.17207Syntology ran 1 of 1 samples · 0 unverified
-
Random-Access Infinite Context Length for Transformers 21 Sep 2023 · 1 repository
-
RCMHA: Relative Convolutional Multi-Head Attention for Natural Language Modelling 7 Aug 2023 · 1 repository · arXiv:2308.03429
-
Landmark Attention: Random-Access Infinite Context Length for Transformers 25 May 2023 · 2 repositories · arXiv:2305.16300Syntology ran 11 of 13 samples · 2 unverified
-
Transformer-based World Models Are Happy With 100k Interactions 13 Mar 2023 · 1 repository · arXiv:2303.07109Syntology ran 16 of 25 samples · 9 unverified
-
GTR-CTRL: Instrument and Genre Conditioning for Guitar-Focused Music Generation with Transformers 10 Feb 2023 · 0 repositories · arXiv:2302.05393
-
An Comparative Analysis of Different Pitch and Metrical Grid Encoding Methods in the Task of Sequential Music Generation 31 Jan 2023 · 0 repositories · arXiv:2301.13383
-
Efficient Sparsely Activated Transformers 31 Aug 2022 · 0 repositories · arXiv:2208.14580
-
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models 13 Aug 2022 · 9 repositories · arXiv:2208.06677Syntology ran 1 of 1 samples · 0 unverified
-
Recurrent Memory Transformer 14 Jul 2022 · 3 repositories · arXiv:2207.06881Syntology ran 6 of 13 samples · 7 unverified · 2 pointer-only (licence)
-
Emotion-Aware Transformer Encoder for Empathetic Dialogue Generation 24 Apr 2022 · 1 repository · arXiv:2204.11320
-
SinTra: Learning an inspiration model from a single multi-track music segment 21 Apr 2022 · 1 repository · arXiv:2204.09917
-
LiteTransformerSearch: Training-free Neural Architecture Search for Efficient Language Models 4 Mar 2022 · 1 repository · arXiv:2203.02094Syntology ran 1 of 5 samples · 4 unverified
-
Reconsidering the Past: Optimizing Hidden States in Language Models 16 Dec 2021 · 0 repositories · arXiv:2112.08653
-
A Comparative Study of Transformers on Word Sense Disambiguation 30 Nov 2021 · 0 repositories · arXiv:2111.15417
-
How much do language models copy from their training data? Evaluating linguistic novelty in text generation using RAVEN 18 Nov 2021 · 0 repositories · arXiv:2111.09509
-
TACCL: Guiding Collective Algorithm Synthesis using Communication Sketches 8 Nov 2021 · 2 repositories · arXiv:2111.04867
-
Language Modelling via Learning to Rank 13 Oct 2021 · 0 repositories · arXiv:2110.06961
Tasks archive 2025-07-28
20 shown of 96 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections