Methods › Sequential › Recurrent Neural Networks › SHA-RNN

Single Headed Attention RNN

SHA-RNN

2 papers tagged archive 2025-07-28

Introduced by Stephen Merity in Single Headed Attention RNN: Stop Thinking With Your Head

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

SHA-RNN, or Single Headed Attention RNN, is a recurrent neural network, and language model when combined with an embedding input and softmax classifier, based on a core LSTM component and a single-headed attention module. Other design choices include a Boom feedforward layer and the use of layer normalization. The guiding principles of the author were to ensure simplicity in the architecture and to keep computational costs bounded (the model was originally trained with a single GPU).

PaperSource

Papers archive 2025-07-28

2 shown of 2, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

4 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
GPU1
Hyperparameter Optimization1
Language Modeling1
Language Modelling1

Usage over time archive 2025-07-28

Papers per year tagged with SHA-RNN: 2019 to 2021, peak 1 1 0 2019: 1 paper 2019 2020: 0 papers 2020 2021: 1 paper 2021
Papers per year the archive tags with this method, by the paper's archive date (2 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Recurrent Neural NetworksLanguage Models

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections