Browse State-of-the-Art › Topic Models
Topic Models
229 papers with code · 6 benchmarks · 17 datasets archive 2025-07-28
A topic model is a type of statistical model for discovering the abstract "topics" that occur in a collection of documents. Topic modeling is a frequently used text-mining tool for the discovery of hidden semantic structures in a text body.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| 20NewsGroups (6 rows) | vONTSS | vONTSS: vMF based semi-supervised neural topic modeling with... | — | — | Compare |
| AG News (6 rows) | vONTSS | vONTSS: vMF based semi-supervised neural topic modeling with... | — | — | Compare |
| 20 Newsgroups (2 rows) | Bayesian SMM | Learning document embeddings along with their uncertainties | code | — | Compare |
| Arxiv HEP-TH citation graph (2 rows) | JoSH | Hierarchical Topic Mining via Joint Spherical Tree and Text Embedding | code | — | Compare |
| AgNews (1 row) | vONTSS | vONTSS: vMF based semi-supervised neural topic modeling with... | — | — | Compare |
| NYT (1 row) | JoSH | Hierarchical Topic Mining via Joint Spherical Tree and Text Embedding | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
17 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 229 papers with code (881 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
8 Jul 2019 12 repositories listed Syntology ran 3 of 13 samples · 10 unverifiedTo this end, we develop the Embedded Topic Model (ETM), a generative model of documents that marries traditional topic models with word embeddings.
-
4 Mar 2017 6 repositories listed Syntology ran 1 of 8 samples · 7 unverified · 1 pointer-only (licence)A promising approach to address this problem is autoencoding variational Bayes (AEVB), but it has proven diffi- cult to apply to topic models in practice.
-
19 Nov 2015 6 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 1 pointer-only (licence)We validate this framework on two very different text modelling applications, generative document modelling and supervised question answering.
-
6 May 2016 5 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedDistributed dense word vectors have been shown to be effective at capturing token-level semantic and syntactic regularities in language, while topic models can form interpretable representations over documents.
-
29 May 2019 4 repositories listedTo address this challenge, we develop causally sufficient embeddings, low-dimensional document representations that preserve sufficient information for causal identification and allow for efficient estimation of causal…
-
11 Mar 2022 3 repositories listed Syntology ran 1 of 8 samples · 7 unverifiedBERTopic generates coherent topics and remains competitive across a variety of benchmarks involving classical models and those that follow the more recent clustering approach of topic modeling.
-
8 Apr 2020 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedTopic models extract groups of words from documents, whose interpretation as a topic hopefully allows for a better understanding of the data.
-
3 Jul 2018 3 repositories listedThis paper describes an alternative approach to discover topics based on Min-Hashing, which can handle massive text corpora and large vocabularies using modest computer hardware and does not require to fix the number of…
-
1 Jul 2017 3 repositories listedUnlike topic models which typically assume independently generated words, word embedding models encourage words that appear in similar contexts to be located close to each other in the embedding space.
-
25 May 2017 3 repositories listed Syntology ran 6 of 6 samples · 0 unverifiedMost real-world document collections involve various types of metadata, such as author, source, and date, and yet the most commonly-used approaches to modeling text corpora ignore this information.
-
13 Jun 2024 2 repositories listedTopic modeling has been a widely used tool for unsupervised text analysis.
-
28 May 2024 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedWe further propose a novel Embedding Transport Plan (ETP) method.
-
28 May 2024 2 repositories listedHowever, existing models suffer from repetitive topic and unassociated topic issues, failing to reveal the evolution and hindering further applications.
-
27 Jan 2024 2 repositories listedIn this paper, we present a comprehensive survey on neural topic models concerning methods, applications, and challenges.
-
7 Jun 2023 2 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 2 pointer-only (licence)Topic models have been prevalent for decades with various applications.
-
7 Apr 2023 2 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 2 pointer-only (licence)Instead of the direct alignment in previous work, we propose a topic alignment with mutual information method.
-
27 Mar 2023 2 repositories listedTopic modeling is a dominant method for exploring document collections on the web and in digital libraries.
-
27 Mar 2023 2 repositories listedTopic modeling has emerged as a dominant method for exploring large document collections.
-
1 Oct 2022 2 repositories listedNeural topic models have been widely used in discovering the latent semantics from a corpus.
-
Principled Analysis of Energy Discourse across Domains with Thesaurus-based Automatic Topic Labeling1 Dec 2021 2 repositories listedWith the increasing impact of Natural Language Processing tools like topic models in social science research, the experimental rigor and comparability of models and datasets has come under scrutiny.
-
25 Oct 2021 2 repositories listed Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)Recent empirical studies show that adversarial topic models (ATM) can successfully capture semantic patterns of the document by differentiating a document with another dissimilar sample.
-
13 Sep 2021 2 repositories listed Syntology ran 2 of 9 samples · 7 unverified · 2 pointer-only (licence)Phrase representations derived from BERT often do not exhibit complex phrasal compositionality, as the model relies instead on lexical similarity to determine semantic relatedness.
-
5 Jul 2021 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedTo address the standardization gap, we systematically evaluate a dominant classical model and two state-of-the-art neural models on two commonly used datasets.
-
19 Feb 2021 2 repositories listedThe first one is a RoBERTa [10] based model built over these abstracts.
-
1 Nov 2020 2 repositories listedTopic models have been prevailing for many years on discovering latent semantics while modeling long documents.
-
19 Aug 2020 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedDistributed representations of documents and words have gained popularity due to their ability to capture semantics of words and documents.
-
16 Apr 2020 2 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedThey all cover the same content, but the linguistic differences make it impossible to use traditional, bag-of-word-based topic models.
-
20 Nov 2019 2 repositories listedThis research proposes a new (old) metric for evaluating goodness of fit in topic models, the coefficient of determination, or R².
-
20 Aug 2019 2 repositories listedWe present Bayesian subspace multinomial model (Bayesian SMM), a generative log-linear model that learns to represent documents in the form of Gaussian distributions, thereby encoding the uncertainty in its co-variance.
-
2 Nov 2018 2 repositories listedRecently, considerable research effort has been devoted to developing deep architectures for topic models to learn topic structures.
Syntology lines on 15 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections