{"url":"/method/hierarchical-softmax","slug":"hierarchical-softmax","name":"Hierarchical Softmax","full_name":"Hierarchical Softmax","full_name_withheld":false,"description_markdown":"**Hierarchical Softmax** is a is an alternative to [softmax](https://paperswithcode.com/method/softmax) that is faster to evaluate: it is $O\\left(\\log{n}\\right)$ time to evaluate compared to $O\\left(n\\right)$ for softmax. It utilises a multi-layer binary tree, where the probability of a word is calculated through the product of probabilities on each edge on the path to that node. See the Figure to the right for an example of where the product calculation would occur for the word \"I'm\".\r\n\r\n(Introduced by Morin and Bengio)\r\n\r\nImage Credit: [Steven Schmatz](https://www.quora.com/profile/Steven-Schmatz)","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":null,"title":null,"url_on_a_paper_host":false},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Output Functions","url":"/methods/category/output-functions","pwc_aliases":[]}],"n_papers_tagged":12,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits","date":"2025-05-15","arxiv_id":"2505.10202","n_code_links":0,"syntology":null},{"paper":null,"title":"Convergence Rates for Softmax Gating Mixture of Experts","date":"2025-03-05","arxiv_id":"2503.03213","n_code_links":0,"syntology":null},{"paper":null,"title":"Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition","date":"2025-01-29","arxiv_id":"2501.17615","n_code_links":0,"syntology":null},{"paper":"/paper/global-hierarchical-neural-networks-using","title":"Global Hierarchical Neural Networks using Hierarchical Softmax","date":"2023-08-02","arxiv_id":"2308.01210","n_code_links":1,"syntology":null},{"paper":"/paper/hierarchical-softmax-for-end-to-end-low","title":"Hierarchical Softmax for End-to-End Low-resource Multilingual Speech Recognition","date":"2022-04-08","arxiv_id":"2204.03855","n_code_links":1,"syntology":null},{"paper":"/paper/probabilistic-label-trees-for-extreme-multi","title":"Probabilistic Label Trees for Extreme Multi-label Classification","date":"2020-09-23","arxiv_id":"2009.11218","n_code_links":2,"syntology":null},{"paper":null,"title":"On the computational complexity of the probabilistic label tree algorithms","date":"2019-06-01","arxiv_id":"1906.00294","n_code_links":0,"syntology":null},{"paper":null,"title":"Effectiveness of Hierarchical Softmax in Large Scale Classification Tasks","date":"2018-12-13","arxiv_id":"1812.05737","n_code_links":0,"syntology":null},{"paper":"/paper/a-no-regret-generalization-of-hierarchical","title":"A no-regret generalization of hierarchical softmax to extreme multi-label classification","date":"2018-10-27","arxiv_id":"1810.11671","n_code_links":1,"syntology":null},{"paper":null,"title":"Self-organized Hierarchical Softmax","date":"2017-07-26","arxiv_id":"1707.08588","n_code_links":0,"syntology":null},{"paper":"/paper/word2vec-parameter-learning-explained","title":"word2vec Parameter Learning Explained","date":"2014-11-11","arxiv_id":"1411.2738","n_code_links":8,"syntology":{"ran":4,"of":9,"unverified":5,"pointer_only":4}},{"paper":"/paper/distributed-representations-of-words-and-1","title":"Distributed Representations of Words and Phrases and their Compositionality","date":"2013-10-16","arxiv_id":"1310.4546","n_code_links":51,"syntology":{"ran":6,"of":26,"unverified":20,"pointer_only":4}}],"papers_shown":12,"tasks":[{"task":"/task/classification-1","name":"Classification","papers":3},{"task":"/task/classification","name":"General Classification","papers":3},{"task":"/task/extreme-multi-label-classification","name":"Extreme Multi-Label Classification","papers":2},{"task":"/task/language-modeling","name":"Language Modeling","papers":2},{"task":"/task/language-modelling","name":"Language Modelling","papers":2},{"task":"/task/multi-label-classification-2","name":"MUlTI-LABEL-ClASSIFICATION","papers":2},{"task":"/task/multi-label-classification","name":"Multi-Label Classification","papers":2},{"task":"/task/speech-recognition","name":"Speech Recognition","papers":2},{"task":"/task/speech-recognition-1","name":"speech-recognition","papers":2},{"task":"/task/automatic-speech-recognition-2","name":"Automatic Speech Recognition","papers":1},{"task":"/task/automatic-speech-recognition","name":"Automatic Speech Recognition (ASR)","papers":1},{"task":"/task/decoder","name":"Decoder","papers":1},{"task":"/task/mixture-of-experts","name":"Mixture-of-Experts","papers":1},{"task":"/task/multi-class-classification","name":"Multi-class Classification","papers":1},{"task":"/task/quantization","name":"Quantization","papers":1},{"task":"/task/sentence","name":"Sentence","papers":1},{"task":"/task/sentence-compression","name":"Sentence Compression","papers":1},{"task":"/task/text-classification","name":"Text Classification","papers":1},{"task":"/task/parameter-estimation","name":"parameter estimation","papers":1},{"task":"/task/text-classification-1","name":"text-classification","papers":1}],"tasks_shown":20,"n_tasks":20,"usage_by_year":[{"year":"2013","papers":1},{"year":"2014","papers":1},{"year":"2017","papers":1},{"year":"2018","papers":2},{"year":"2019","papers":1},{"year":"2020","papers":1},{"year":"2022","papers":1},{"year":"2023","papers":1},{"year":"2025","papers":3}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/hierarchical-softmax"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}