Datasets › CPED

CPED (Chinese Personalized and Emotional Dialogue)

Introduced by YiRong Chen et al. in CPED: A Large-Scale Chinese Personalized and Emotional Dialogue Dataset for Conversational AI29 May 2022 archive 2025-07-28

We construct a dataset named CPED from 40 Chinese TV shows. CPED consists of multisource knowledge related to empathy and personal characteristic. This knowledge covers 13 emotions, gender, Big Five personality traits, 19 dialogue acts and other knowledge.

  • We build a multiturn Chinese Personalized and Emotional Dialogue dataset called CPED. To the best of our knowledge, CPED is the first Chinese personalized and emotional dialogue dataset. CPED contains 12K dialogues and 133K utterances with multi-modal context. Therefore, it can be used in both complicated dialogue understanding and human-like conversation generation.
  • CPED has been annotated with 3 character attributes (name, gender age), Big Five personality traits, 2 types of dynamic emotional information (sentiment and emotion) and DAs. The personality traits and emotions can be used as prior external knowledge for open-domain conversation generation, making the conversation system have a good command of personification capabilities.
  • We propose three tasks for CPED: personality recognition in conversations (PRC), emotion recognition in conversations (ERC), and personalized and emotional conversation (PEC). A set of experiments verify the importance of using personalities and emotions as prior external knowledge for conversation generation.

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

11 shown of 11 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 15. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
CPED: A Large-Scale Chinese Personalized and Emotional Dialogue Dataset for Conversational AI 1 14 29 May 2022 not harvested
EmoBERTa: Speaker-Aware Emotion Recognition in Conversation with RoBERTa 1 1 26 Aug 2021 not harvested
DialogXL: All-in-One XLNet for Multi-Party Conversation Emotion Recognition 4 1 16 Dec 2020 not harvested
DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation 2 1 30 Aug 2019 ran 0 of 4 samples (4 unverified)
DialogueRNN: An Attentive RNN for Emotion Detection in Conversations 2 1 1 Nov 2018 not harvested
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 534 1 11 Oct 2018 ran 204 of 659 samples (455 unverified; 149 pointer-only for licence)
Context-Dependent Sentiment Analysis in User-Generated Videos 2 1 1 Jul 2017 not harvested
Bag of Tricks for Efficient Text Classification 65 1 6 Jul 2016 ran 2 of 9 samples (7 unverified; 2 pointer-only for licence)
Recurrent Neural Network for Text Classification with Multi-Task Learning 0 1 17 May 2016 not harvested
Recurrent Convolutional Neural Networks for Text Classification 1 1 1 Jan 2015 not harvested
Convolutional Neural Networks for Sentence Classification 118 1 25 Aug 2014 ran 19 of 77 samples (58 unverified; 15 pointer-only for licence)

Dataset loaders archive 2025-07-28

scutcyr/CPEDpytorch

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Apache-2.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • CPED

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections