Home › Datasets › task › Intent Detection

Intent Detection datasets

archive 2025-07-28

25 datasets carry the task tag "Intent Detection" (the task itself: Intent Detection), ordered by the archive's paper count. Page 1 of 1: 25 shown of 25. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Intent Detection datasets 1–25 of 25

MultiWOZ (Multi-domain Wizard-of-Oz)
The Multi-domain Wizard-of-Oz (MultiWOZ) dataset is a large-scale human-human conversational corpus spanning over seven domains, containing 8438 multi-turn dialogues, with each dialogue averaging 14 turns.
328 papers · 8 benchmarks
SNIPS (SNIPS Natural Language Understanding benchmark)
The SNIPS Natural Language Understanding benchmark is a dataset of over 16,000 crowdsourced queries distributed among 7 user intents of various complexity: SearchCreativeWork (e.g.
256 papers · 6 benchmarks
This dataset is for evaluating the performance of intent classification systems in the presence of "out-of-scope" queries, i.e., queries that do not fall into any of the system-supported intent classes.
87 papers · 5 benchmarks
This project contains natural language data for human-robot interaction in home domain which we collected and annotated for evaluating NLU Services/platforms.
66 papers · 3 benchmarks
The Dialog State Tracking Challenges 2 & 3 (DSTC2&3) were research challenge focused on improving the state of the art in tracking the state of spoken dialog systems.
33 papers · 5 benchmarks
Dataset is constructed from single intent dataset ATIS.
28 papers · 2 benchmarks
Dataset is constructed from single intent dataset SNIPS.
26 papers · 2 benchmarks
SVIRO (Synthetic Vehicle Interior Rear Seat Occupancy Dataset)
Contains bounding boxes for object detection, instance segmentation masks, keypoints for pose estimation and depth images for each synthetic scenery as well as images for each individual seat for classification.
15 papers · 0 benchmarks
HINT3 is a dataset for intent detection.
13 papers · 0 benchmarks
SPEECH-COCO contains speech captions that are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images.
9 papers · 0 benchmarks
A dataset with a single banking domain, includes both general Out-of-Scope (OOD-OOS) queries and In-Domain but Out-of-Scope (ID-OOS) queries, where ID-OOS queries are semantically similar intents/queries with in-scope intents.
6 papers · 1 benchmark
CAIS (Chinese Artificial Intelligence Speakers)
We collect utterances from the Chinese Artificial Intelligence Speakers (CAIS), and annotate them with slot tags and intent labels.
6 papers · 2 benchmarks
ATIS (vi) (Vietnamese Intent Detection and Slot Filling)
This is a dataset for intent detection and slot filling for the Vietnamese language.
4 papers · 2 benchmarks
ITALIC: An ITALian Intent Classification Dataset ITALIC is an intent classification dataset for the Italian language, which is the first of its kind.
4 papers · 0 benchmarks
A dataset with two separate domains, i.e., the "Banking'' domain and the "Credit cards'' domain with both general Out-of-Scope (OOD-OOS) queries and In-Domain but Out-of-Scope (ID-OOS) queries, where ID-OOS queries are semantically similar…
3 papers · 0 benchmarks
ClarQ, consists of ∼2M examples distributed across 173 domains of stackexchange.
3 papers · 0 benchmarks
MDID (Multimodal Document Intent Dataset)
The Multimodal Document Intent Dataset (MDID) is a dataset for computing author intent from multimodal data from Instagram.
3 papers · 0 benchmarks
Almawave-SLU is the first Italian dataset for Spoken Language Understanding (SLU).
2 papers · 0 benchmarks
DialogUSR dataset covers 23 domains with a multi-step crowd-sourcing procedure.
2 papers · 0 benchmarks
The PATIS is a Persian language dataset for intent detection and slot filling.
2 papers · 2 benchmarks
ProSLU (Profile-based Spoken Language Understanding)
In the paper, to bridge the research gap, we propose a new and important task, Profile-based Spoken Language Understanding (ProSLU), which requires a model not only depends on the text but also on the given supporting profile information.
2 papers · 2 benchmarks
MIPD (Manipulation and Intention In a Novel Corpus of Polish Disinformation)
A novel corpus of 15,356 Polish web articles, including articles identified as containing disinformation.
1 paper · 0 benchmarks
NoMusic - The Norwegian Multi-Dialectal Slot and Intent Detection Corpus https://aclanthology.org/2024.vardial-1.9/
1 paper · 0 benchmarks
diaforge-utc-r-0725 (DiaFORGE UTC: Unified Tool-Calling Conversations Dataset)
Dataset for our paper Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky which includes 5000 enterprise tools and the corresponding dialogues generated using DiaFORGE UTC data engine.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.