Browse State-of-the-Art › Text Clustering
Text Clustering
38 papers with code · 3 benchmarks · 5 datasets archive 2025-07-28
Grouping a set of texts in such a way that objects in the same group (called a cluster) are more similar (in some sense) to each other than to those in other groups (clusters). (Source: Adapted from Wikipedia)
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| MTEB (31 rows) | ST5-XXL | MTEB: Massive Text Embedding Benchmark | code | Syntology ran 3 of 13 samples · 10 unverified | Compare |
| 20 Newsgroups (2 rows) | G-BAT | Neural Topic Modeling with Bidirectional Adversarial Training | code | — | Compare |
| Urdu News Headlines Dataset (1 row) | Vector Space Model | Clustering Urdu News Using Headlines | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
5 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
3 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 38 papers with code (123 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
13 Oct 2022 5 repositories listed Syntology ran 3 of 13 samples · 10 unverifiedMTEB spans 8 embedding tasks covering a total of 58 datasets and 112 languages.
-
1 Jun 2015 3 repositories listed
-
16 Dec 2021 2 repositories listedText clustering methods were traditionally incorporated into multi-document summarization (MDS) as a means for coping with considerable information repetition.
-
24 Mar 2021 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedUnsupervised clustering aims at discovering the semantic categories of data according to some distance measured in the representation space.
-
16 Dec 2020 2 repositories listedIn this work, we propose an effective method, Deep Aligned Clustering, to discover new intents with the aid of the limited known intent data.
-
15 Jun 2020 2 repositories listedThe dissimilarity mixture autoencoder (DMAE) is a neural network model for feature-based clustering that incorporates a flexible dissimilarity function and can be integrated into any kind of deep learning architecture.
-
1 May 2025 1 repository listedAs a fundamental task in Information Retrieval and Computational Linguistics, sentence representation has profound implications for a wide range of practical applications such as text clustering, content analysis,…
-
25 Jan 2025 1 repository listedBy solving this OT problem, we can yield reliable pseudo-labels that simultaneously account for sample-to-sample semantic consistency and sample-to-cluster global structure information.
-
7 Jan 2025 1 repository listedThen, the former module uses the more discriminative consistent representation to produce reliable supervision information for assist clustering, while the latter module explores similarity relationships and consistent…
-
30 Sep 2024 1 repository listed Syntology ran 6 of 6 samples · 0 unverified · 6 pointer-only (licence)Second, after integrating similar labels generated by the LLM, we prompt the LLM to assign the most appropriate label to each sample in the dataset.
-
23 Aug 2024 1 repository listedInterpretable clustering algorithms aim to group similar data points while explaining the obtained groups to support knowledge discovery and pattern recognition tasks.
-
1 Aug 2024 1 repository listedSocial media are a critical component of the information ecosystem during public health crises.
-
12 May 2024 1 repository listedThe generative LLM shows good agreement with the human reviewers, and is suggested as a means to bridge the `validation gap' which often exists between cluster production and cluster interpretation.
-
20 Feb 2024 1 repository listedThis paper explores an empirical approach to learn more discriminantive sentence representations in an unsupervised fashion.
-
2 Jul 2023 1 repository listed Syntology ran 0 of 6 samples · 6 unverifiedIn this paper, we ask whether a large language model can amplify an expert's guidance to enable query-efficient, few-shot semi-supervised text clustering.
-
24 May 2023 1 repository listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)First, we prompt ChatGPT for insights on clustering perspective by constructing hard triplet questions <does A better correspond to B than C>, where A, B and C are similar data points that belong to different clusters…
-
23 May 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedTo tackle the above issues, we propose a Robust Short Text Clustering (RSTC) model to improve robustness against imbalanced and noisy data.
-
4 May 2023 1 repository listedFor example, a three star rating (out of five) may be incongruous with the review text, which may be more suitable for a five star review.
-
2 Mar 2023 1 repository listedIn this work, we propose DeepLens, an interactive system that helps users detect and explore OOD issues in massive text corpora.
-
19 Dec 2022 1 repository listedText data mining is the process of deriving essential information from language text.
-
26 Jul 2022 1 repository listedOur sentence encoder can be trained in less than a day on a single graphics card, achieving high performance on a diverse set of sentence-level tasks.
-
1 Jun 2022 1 repository listedIn this paper we describe an experiment for the application of text clustering techniques to dossiers of amendments to proposed legislation discussed in the Italian Senate.
-
9 May 2022 1 repository listed Syntology ran 1 of 5 samples · 4 unverifiedWe present EASE, a novel method for learning sentence embeddings via contrastive learning between sentences and their related entities.
-
1 Feb 2022 1 repository listedWe first extend the concept of subspace clustering to co-clustering, which has been extensively used on document-term matrices due to the resulting interplay between the document and term representations.
-
16 Jan 2022 1 repository listedText clustering methods were traditionally incorporated into multi-document summarization (MDS) as a means for coping with considerable information repetition.
-
1 Nov 2021 1 repository listedA reliable clustering algorithm for task-oriented dialogues can help developer analysis and define dialogue tasks efficiently.
-
16 Sep 2021 1 repository listedHere we analyze the sentence representations learned by NMT Transformers and show that these explicitly include the information on text domains, even after only seeing the input sentences without domains labels.
-
1 Aug 2021 1 repository listedExisting supervised models for text clustering find it difficult to directly optimize for clustering results.
-
30 Jul 2021 1 repository listedSpherical k-Means is frequently used to cluster document collections because it performs reasonably well in many settings and is computationally efficient.
-
11 Oct 2020 1 repository listedTopic detection is the task of determining and tracking hot topics in social media.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections