Browse State-of-the-Art › Short Text Clustering
Short Text Clustering
18 papers with code · 8 benchmarks · 2 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
8 leaderboard tables shown for this task, 8 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Stackoverflow (5 rows) | TCL | Twin Contrastive Learning for Online Clustering | code | — | Compare |
| Biomedical (4 rows) | TCL | Twin Contrastive Learning for Online Clustering | code | — | Compare |
| Searchsnippets (4 rows) | SCCL | Supporting Clustering with Contrastive Learning | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| AG News (1 row) | SCCL | Supporting Clustering with Contrastive Learning | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| GoogleNews-S (1 row) | SCCL | Supporting Clustering with Contrastive Learning | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| GoogleNews-T (1 row) | SCCL | Supporting Clustering with Contrastive Learning | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| GoogleNews-TS (1 row) | SCCL | Supporting Clustering with Contrastive Learning | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| Tweet (1 row) | SCCL | Supporting Clustering with Contrastive Learning | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
18 shown of 18 papers with code (34 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
1 Jun 2015 3 repositories listed
-
21 Oct 2022 2 repositories listedSpecifically, we find that when the data is projected into a feature space with a dimensionality of the target cluster number, the rows and columns of its feature matrix correspond to the instance and cluster…
-
24 Mar 2021 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedUnsupervised clustering aims at discovering the semantic categories of data according to some distance measured in the representation space.
-
16 Dec 2020 2 repositories listedIn this work, we propose an effective method, Deep Aligned Clustering, to discover new intents with the aid of the limited known intent data.
-
25 Jan 2025 1 repository listedBy solving this OT problem, we can yield reliable pseudo-labels that simultaneously account for sample-to-sample semantic consistency and sample-to-cluster global structure information.
-
7 Jan 2025 1 repository listedThen, the former module uses the more discriminative consistent representation to produce reliable supervision information for assist clustering, while the latter module explores similarity relationships and consistent…
-
1 Aug 2024 1 repository listedSocial media are a critical component of the information ecosystem during public health crises.
-
12 May 2024 1 repository listedThe generative LLM shows good agreement with the human reviewers, and is suggested as a means to bridge the `validation gap' which often exists between cluster production and cluster interpretation.
-
23 May 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedTo tackle the above issues, we propose a Robust Short Text Clustering (RSTC) model to improve robustness against imbalanced and noisy data.
-
9 May 2022 1 repository listed Syntology ran 1 of 5 samples · 4 unverifiedWe present EASE, a novel method for learning sentence embeddings via contrastive learning between sentences and their related entities.
-
1 Aug 2021 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)This paper develops the DECAF algorithm that addresses these challenges by learning models enriched by label metadata that jointly learn model parameters and feature representations using deep networks and offer…
-
31 Jul 2021 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedThis paper presents ECLARE, a scalable deep learning architecture that incorporates not only label text, but also label correlations, to offer accurate real-time predictions within a few milliseconds.
-
30 Jul 2021 1 repository listedSpherical k-Means is frequently used to cluster document collections because it performs reasonably well in many settings and is computationally efficient.
-
22 May 2020 1 repository listedIn this paper, we present an intent discovery framework that involves 4 primary steps: Extraction of textual utterances from a conversation using a pre-trained domain agnostic Dialog Act Classifier (Data Extraction),…
-
31 Jan 2020 1 repository listedShort text clustering is a challenging task due to the lack of signal contained in such short texts.
-
20 Nov 2019 1 repository listedIdentifying new user intents is an essential task in the dialogue system.
-
1 Aug 2019 1 repository listedShort text clustering is a challenging problem when adopting traditional bag-of-words or TF-IDF representations, since these lead to sparse vector representations of the short texts.
-
1 Jan 2017 1 repository listedShort text clustering is a challenging problem due to its sparseness of text representation.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections