Home › Datasets › task › Dialogue Generation

Dialogue Generation datasets

archive 2025-07-28

34 datasets carry the task tag "Dialogue Generation" (the task itself: Dialogue Generation), ordered by the archive's paper count. Page 1 of 1: 34 shown of 34. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Dialogue Generation datasets 1–34 of 34

The Reddit dataset is a graph dataset from Reddit posts made in the month of September, 2014.
699 papers · 8 benchmarks
OpenSubtitles is collection of multilingual parallel corpora.
214 papers · 3 benchmarks
A new open-vocabulary language modelling benchmark derived from books.
147 papers · 1 benchmark
DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 dialogues with corresponding manually labeled summaries and topics.
62 papers · 2 benchmarks
OpenDialKG contains utterance from 15K human-to-human role-playing dialogs is manually annotated with ground-truth reference to corresponding entities and paths from a large-scale KG with 1M+ facts.
55 papers · 0 benchmarks
UDC (Ubuntu Dialogue Corpus)
Ubuntu Dialogue Corpus (UDC) is a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words.
46 papers · 8 benchmarks
Doc2Dial (Doc2Dial: Document-grounded Dialogue)
For goal-oriented document-grounded dialogs, it often involves complex contexts for identifying the most relevant information, which requires better understanding of the inter-relations between conversations and documents.
36 papers · 0 benchmarks
MultiDoc2Dial (MultiDoc2Dial: Modeling Dialogues Grounded in Multiple Documents)
MultiDoc2Dial is a new task and dataset on modeling goal-oriented dialogues grounded in multiple documents.
26 papers · 0 benchmarks
LCCC (Large-scale Cleaned Chinese Conversation corpus)
Contains a base version (6.8million dialogues) and a large version (12.0 million dialogues).
17 papers · 0 benchmarks
SODA is a high-quality social dialogue dataset.
17 papers · 0 benchmarks
FaithDial is a new benchmark for hallucination-free dialogues, by editing hallucinated responses in the Wizard of Wikipedia (WoW) benchmark.
16 papers · 0 benchmarks
CMU DoG (CMU Document Grounded Conversations Dataset)
This is a document grounded dataset for text conversations.
15 papers · 0 benchmarks
CPED (Chinese Personalized and Emotional Dialogue)
We construct a dataset named CPED from 40 Chinese TV shows.
15 papers · 3 benchmarks
PersonalDialog is a large-scale multi-turn dialogue dataset containing various traits from a large number of speakers.
13 papers · 0 benchmarks
Most existing dialogue systems fail to respond properly to potentially unsafe user utterances by either ignoring or passively agreeing with them.
13 papers · 1 benchmark
KaMed is a knowledge-aware medical dialogue dataset, which contains over 60,000 medical dialogue sessions with 5,682 entities (such as Asthma and Atropine).
9 papers · 0 benchmarks
OpenViDial is a large-scale open-domain dialogue dataset with visual contexts.
8 papers · 0 benchmarks
FusedChat is an inter-mode dialogue dataset.
6 papers · 1 benchmark
OTTers is a dataset of human one-turn topic transitions.
6 papers · 0 benchmarks
- A large scale Chinese multi-modal dialogue corpus (120.84K dialogues and 198.82 K images).
5 papers · 0 benchmarks
MetaLWOz (Meta-Learning Wizard-of-Oz)
Collected by leveraging background knowledge from a larger, more highly represented dialogue source.
4 papers · 0 benchmarks
The high-quality multi-turn dialogue dataset, which has a total of 3,134 multi-turn consultation dialogues.
3 papers · 0 benchmarks
MDIA is a large-scale multilingual benchmark for dialogue generation.
3 papers · 0 benchmarks
WDC-Dialogue is a dataset built from the Chinese social media to train EVA.
3 papers · 0 benchmarks
The general multi-turn dialogue evaluation dataset with nine topics.
2 papers · 0 benchmarks
CareCall (CareCall for Seniors)
carecall is a Korean dialogue dataset for role-satisfying dialogue systems.
2 papers · 0 benchmarks
OpenViDial 2.0 is a larger-scale open-domain multi-modal dialogue dataset compared to the previous version OpenViDial 1.0.
2 papers · 1 benchmark
Arabic-ToD (Arabic-ToD: Arabic Task Oriented Dialogue dataset)
The Arabic-TOD dataset is based on the BiToD dataset.
1 paper · 0 benchmarks
Harry Potter Dialogue is the first dialogue dataset that integrates with scene, attributes and relations which are dynamically changed as the storyline goes on.
1 paper · 2 benchmarks
JDDC 2.0 is a large-scale multimodal multi-turn dialogue dataset collected from a mainstream Chinese E-commerce platform JD.com, containing about 246 thousand dialogue sessions, 3 million utterances, and 507 thousand images, along with…
1 paper · 0 benchmarks
To evaluate our proposed strategy of asynchronous communication for LLMs, we run games of Mafia with human players, incorporating an LLM-based agent as an additional player, within an asynchronous chat environment.
1 paper · 0 benchmarks
MultiRefKGC (multi-reference KGC)
MultiRefKGC is a dataset created from conversations from Reddit designed for Knowledge-Grounded Dialogue Generation tasks.
1 paper · 0 benchmarks
> The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents > > Xing Han Lu, Siva Reddy, Harm de Vries > > EACL 2023 | | | | | | | :--: | :--: | :--: | :--: | :--: | | Code | Huggingface | Request on…
1 paper · 1 benchmark
diaforge-utc-r-0725 (DiaFORGE UTC: Unified Tool-Calling Conversations Dataset)
Dataset for our paper Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky which includes 5000 enterprise tools and the corresponding dialogues generated using DiaFORGE UTC data engine.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.