{"url":"/method/sha-rnn","slug":"sha-rnn","name":"SHA-RNN","full_name":"Single Headed Attention RNN","full_name_withheld":false,"description_markdown":"**SHA-RNN**, or **Single Headed Attention RNN**, is a recurrent neural network, and language model when combined with an embedding input and [softmax](https://paperswithcode.com/method/softmax) classifier, based on a core [LSTM](https://paperswithcode.com/method/lstm) component and a [single-headed attention](https://paperswithcode.com/method/single-headed-attention) module. Other design choices include a Boom feedforward layer and the use of [layer normalization](https://paperswithcode.com/method/layer-normalization). The guiding principles of the author were to ensure simplicity in the architecture and to keep computational costs bounded (the model was originally trained with a single GPU).","description_state":"present","introduced_year":null,"introduced_by":{"title":"Single Headed Attention RNN: Stop Thinking With Your Head","paper":"/paper/single-headed-attention-rnn-stop-thinking","first_author":"Stephen Merity","n_authors":1,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/single-headed-attention-rnn-stop-thinking"},"source":{"url":"https://arxiv.org/abs/1911.11423v2","title":"Single Headed Attention RNN: Stop Thinking With Your Head","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Sequential","area_id":"sequential","collection":"Recurrent Neural Networks","url":"/methods/category/recurrent-neural-networks","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Language Models","url":"/methods/category/language-models","pwc_aliases":[]}],"n_papers_tagged":2,"archive_num_papers":2,"papers_newest_first":[{"paper":null,"title":"SHAQ: Single Headed Attention with Quasi-Recurrence","date":"2021-08-18","arxiv_id":"2108.08207","n_code_links":0,"syntology":null},{"paper":"/paper/single-headed-attention-rnn-stop-thinking","title":"Single Headed Attention RNN: Stop Thinking With Your Head","date":"2019-11-26","arxiv_id":"1911.11423","n_code_links":5,"syntology":null}],"papers_shown":2,"tasks":[{"task":null,"name":"GPU","papers":1},{"task":"/task/hyperparameter-optimization","name":"Hyperparameter Optimization","papers":1},{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1}],"tasks_shown":4,"n_tasks":4,"usage_by_year":[{"year":"2019","papers":1},{"year":"2021","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/sha-rnn"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}