Papers › EmoWOZ: A Large-Scale Corpus and Labelling Scheme for Emotion Recognition in...

EmoWOZ: A Large-Scale Corpus and Labelling Scheme for Emotion Recognition in Task-Oriented Dialogue Systems

10 Sep 2021LREC 2022 6arXiv:2109.04919archive 2025-07-28

Shutong Feng, Nurul Lubis, Christian Geishauser, Hsien-Chin Lin, Michael Heck, Carel van Niekerk, Milica Gašić

The ability to recognise emotions lends a conversational artificial intelligence a human touch. While emotions in chit-chat dialogues have received substantial attention, emotions in task-oriented dialogues remain largely unaddressed. This is despite emotions and dialogue success having equally important roles in a natural system. Existing emotion-annotated task-oriented corpora are limited in size, label richness, and public availability, creating a bottleneck for downstream tasks. To lay a foundation for studies on emotions in task-oriented dialogues, we introduce EmoWOZ, a large-scale manually emotion-annotated corpus of task-oriented dialogues. EmoWOZ is based on MultiWOZ, a multi-domain task-oriented dialogue dataset. It contains more than 11K dialogues with more than 83K emotion annotations of user utterances. In addition to Wizard-of-Oz dialogues from MultiWOZ, we collect human-machine dialogues within the same set of domains to sufficiently cover the space of various emotions that can happen during the lifetime of a data-driven dialogue system. To the best of our knowledge, this is the first large-scale open-source corpus of its kind. We propose a novel emotion labelling scheme, which is tailored to task-oriented dialogues. We report a set of experimental results to show the usability of this corpus for emotion recognition and state tracking in task-oriented dialogues.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Emotion RecognitionEmotion Recognition in ConversationTask-Oriented Dialogue Systems

Datasets

Introduced by this paper, per the archive.

EmoWOZ

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Emotion Recognition in Conversation EmoWoz COSMIC Macro F1 61.12 #1 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz COSMIC Macro F1 (w/o Neutral) 56.34 #1 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz COSMIC Weighted F1 85.94 #1 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz COSMIC Weighted F1 (w/o Neutral) 77.09 #1 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz ContextBERT Macro F1 59.79 #2 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz ContextBERT Macro F1 (w/o Neutral) 54.30 #2 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz ContextBERT Weighted F1 88.33 #2 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz ContextBERT Weighted F1 (w/o Neutral) 79.67 #2 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz DialogueRNN-BERT Macro F1 57.10 #3 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz DialogueRNN-BERT Macro F1 (w/o Neutral) 52.15 #3 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz DialogueRNN-BERT Weighted F1 83.41 #3 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz DialogueRNN-BERT Weighted F1 (w/o Neutral) 75.50 #3 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz BERT Macro F1 55.80 #4 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz BERT Macro F1 (w/o Neutral) 50.14 #4 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz BERT Weighted F1 84.83 #4 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz BERT Weighted F1 (w/o Neutral) 73.55 #4 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz DialogueRNN-GloVe Macro F1 46.33 #5 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz DialogueRNN-GloVe Macro F1 (w/o Neutral) 40.14 #5 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz DialogueRNN-GloVe Weighted F1 80.76 #5 of 6 Archive leaderboard report
Emotion Recognition in Conversation EmoWoz DialogueRNN-GloVe Weighted F1 (w/o Neutral) 74.56 #5 of 6 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections