Datasets › MultiWOZ

MultiWOZ (Multi-domain Wizard-of-Oz)

Introduced by Paweł Budzianowski et al. in MultiWOZ -- A Large-Scale Multi-Domain Wizard-of-Oz Dataset for Task-Oriented Dialogue Modelling1 Jan 2018 archive 2025-07-28

The Multi-domain Wizard-of-Oz (MultiWOZ) dataset is a large-scale human-human conversational corpus spanning over seven domains, containing 8438 multi-turn dialogues, with each dialogue averaging 14 turns. Different from existing standard datasets like WOZ and DSTC2, which contain less than 10 slots and only a few hundred values, MultiWOZ has 30 (domain, slot) pairs and over 4,500 possible values. The dialogues span seven domains: restaurant, hotel, attraction, taxi, train, hospital and police.

Source: Contents

Image Source: Zhang et al

Benchmarks archive 2025-07-28

All 8 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

15 shown of 15 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 328. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
MIDAS: Multi-level Intent, Domain, And Slot Knowledge Distillation for Multi-turn NLU 1 2 15 Aug 2024 not harvested
TextBox 2.0: A Text Generation Library with Pre-trained Language Models 1 1 26 Dec 2022 ran 1 of 1 samples (0 unverified)
DeepStruct: Pretraining of Language Models for Structure Prediction 1 2 21 May 2022 ran 7 of 13 samples (6 unverified)
GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-Supervised Learning and Explicit Policy Injection 1 2 29 Nov 2021 not harvested
Dialogue State Tracking with a Language Model using Schema-Driven Prompting 1 4 15 Sep 2021 not harvested
Pretraining the Noisy Channel Model for Task-Oriented Dialogue 0 1 18 Mar 2021 not harvested
AuGPT: Auxiliary Tasks and Data Augmentation for End-To-End Dialogue with Pre-Trained Language Models 1 2 9 Feb 2021 not harvested
A Probabilistic End-To-End Task-Oriented Dialog Model with Latent Belief States towards Semi-Supervised Learning 1 1 17 Sep 2020 ran 4 of 10 samples (6 unverified)
Text-to-Text Pre-Training for Data-to-Text Tasks 2 1 21 May 2020 not harvested
SOLOIST: Building Task Bots at Scale with Transfer Learning and Machine Teaching 1 1 11 May 2020 ran 1 of 3 samples (2 unverified)
A Simple Language Model for Task-Oriented Dialogue 1 2 2 May 2020 ran 3 of 14 samples (11 unverified)
Template Guided Text Generation for Task-Oriented Dialogue 0 2 30 Apr 2020 not harvested
Few-shot Natural Language Generation for Task-Oriented Dialog 2 1 27 Feb 2020 not harvested
Task-Oriented Dialog Systems that Consider Multiple Appropriate Responses under the Same Context 6 1 24 Nov 2019 not harvested
Semantically Conditioned Dialog Response Generation via Hierarchical Disentangled Self-Attention 2 1 30 May 2019 ran 2 of 3 samples (1 unverified; 2 pointer-only for licence)

Dataset loaders archive 2025-07-28

4 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

MIT

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • MULTIWOZ 2.0
  • MULTIWOZ 2.1
  • MultiWOZ
  • MULTIWOZ 2.2
  • MULTIWOZ 2.3
  • MULTIWOZ 2.4

6 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections