Browse State-of-the-Art › Intent Discovery
Intent Discovery
20 papers with code · 3 benchmarks · 3 datasets archive 2025-07-28
Given a set of labelled and unlabelled utterances, the idea is to identify existing (known) intents and potential (new intents) intents. This method can be utilised in conversational system setting.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| ATIS (1 row) | k-PCA + HDBSCAN | A Hybrid Architecture for Out of Domain Intent Detection and... | code | — | Compare |
| Persian-ATIS (1 row) | k-PCA + HDBSCAN | A Hybrid Architecture for Out of Domain Intent Detection and... | code | — | Compare |
| SNIPS (1 row) | k-PCA + HDBSCAN | A Hybrid Architecture for Out of Domain Intent Detection and... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
20 shown of 20 papers with code (42 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
15 Aug 2022 2 repositories listedIn our evaluation, we first analyze the quality of the model after adaptive fine-tuning on known classes.
-
13 Sep 2021 2 repositories listedIt is composed of two main modules: open intent detection and open intent discovery.
-
10 Jun 2025 1 repository listedIntent detection aims to identify user intents from natural language inputs, where supervised methods rely heavily on labeled in-domain (IND) data and struggle with out-of-domain (OOD) intents, limiting their practical…
-
31 Mar 2025 1 repository listedNew Intent Discovery (NID) is a crucial task that aims to identify these novel intents while maintaining the capability to recognize existing ones.
-
26 Oct 2024 1 repository listedTo enable better knowledge transfer, we design a prototype learning method integrating the supervised and pseudo signals from IND and OOD samples.
-
12 Jul 2024 1 repository listedThis paper presents a novel and comprehensive solution to enhance both the robustness and efficiency of question answering (QA) systems through supervised contrastive learning (SCL).
-
24 Oct 2023 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedNew Intent Discovery (NID) aims to recognize both new and known intents from unlabeled data with the aid of limited labeled data containing only known intents.
-
16 Oct 2023 1 repository listedIn a practical dialogue system, users may input out-of-domain (OOD) queries.
-
16 Oct 2023 1 repository listedThe tasks of out-of-domain (OOD) intent discovery and generalized intent discovery (GID) aim to extend a closed intent classifier to open-world intent sets, which is crucial to task-oriented dialogue (TOD) systems.
-
31 May 2023 1 repository listedIntent discovery is the task of inferring latent intents from a set of unlabeled utterances, and is a useful step towards the efficient creation of new conversational agents.
-
28 May 2023 1 repository listedPrevious methods suffer from a coupling of pseudo label disambiguation and representation learning, that is, the reliability of pseudo labels relies on representation learning, and representation learning is restricted…
-
16 Apr 2023 1 repository listedNew intent discovery is of great value to natural language processing, allowing for a better understanding of user needs and providing friendly services.
-
7 Mar 2023 1 repository listedOn the other side, a labeled dataset is needed to train a model for Intent Detection in task-oriented dialogue systems.
-
17 Oct 2022 1 repository listedFor OOD clustering stage, we propose a KCC method to form compact clusters by mining true hard negative samples, which bridges the gap between clustering and representation learning.
-
13 Sep 2022 1 repository listedTraditional intent classification models are based on a pre-defined intent set and only recognize limited in-domain (IND) intent classes.
-
25 May 2022 1 repository listedExisting approaches typically rely on a large amount of labeled utterances and employ pseudo-labeling methods for representation learning and clustering, which are label-intensive, inefficient, and inaccurate.
-
24 May 2022 1 repository listedWe use this framework to report baseline intent discovery results over VIRADialogs, that highlight the difficulty of this task.
-
1 May 2022 1 repository listedDiscovering Out-of-Domain(OOD) intents is essential for developing new skills in a task-oriented dialogue system.
-
25 Apr 2021 1 repository listedThis paper presents an unsupervised two-stage approach to discover intents and generate meaningful intent labels automatically from a collection of unlabeled utterances in a domain.
-
22 May 2020 1 repository listedIn this paper, we present an intent discovery framework that involves 4 primary steps: Extraction of textual utterances from a conversation using a pre-trained domain agnostic Dialog Act Classifier (Data Extraction),…
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections