{"url":"/dataset/cped","name":"CPED","full_name":"Chinese Personalized and Emotional Dialogue","description_markdown":"We construct a dataset named CPED from 40 Chinese TV shows. CPED consists of multisource knowledge related to empathy and personal characteristic. This knowledge covers 13 emotions, gender, Big Five personality traits, 19 dialogue acts and other knowledge. \r\n\r\n\r\n* We build a multiturn Chinese Personalized and Emotional Dialogue dataset called CPED. To the best of our knowledge, CPED is the first Chinese personalized and emotional dialogue dataset. CPED contains 12K dialogues and 133K utterances with multi-modal context. Therefore, it can be used in both complicated dialogue understanding and human-like conversation generation.\r\n* CPED has been annotated with 3 character attributes (name, gender age), Big Five personality traits, 2 types of dynamic emotional information (sentiment and emotion) and DAs. The personality traits and emotions can be used as prior external knowledge for open-domain conversation generation, making the conversation system have a good command of personification capabilities.\r\n* We propose three tasks for CPED: personality recognition in conversations (**PRC**), emotion recognition in conversations (**ERC**), and personalized and emotional conversation (**PEC**). A set of experiments verify the importance of using personalities and emotions as prior external knowledge for conversation generation.","description_withheld":null,"homepage":"https://github.com/scutcyr/CPED","introduced_date":"2022-05-29","introduced_date_note":null,"introduced_by":{"paper":"/paper/cped-a-large-scale-chinese-personalized-and-1","title":"CPED: A Large-Scale Chinese Personalized and Emotional Dialogue Dataset for Conversational AI","first_author":"YiRong Chen","url":null},"license":{"name":"Apache-2.0","url":"https://github.com/scutcyr/CPED/blob/main/LICENSE"},"modalities":[{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"},{"name":"Audio","url":"/datasets/modality/audio"}],"tasks":[{"name":"Emotion Recognition","url":"/task/emotion-recognition","datasets_with_task":"/datasets/task/emotion-recognition"},{"name":"Emotion Recognition in Conversation","url":"/task/emotion-recognition-in-conversation","datasets_with_task":"/datasets/task/emotion-recognition-in-conversation"},{"name":"Dialogue Generation","url":"/task/dialogue-generation","datasets_with_task":"/datasets/task/dialogue-generation"},{"name":"Multimodal Emotion Recognition","url":"/task/multimodal-emotion-recognition","datasets_with_task":"/datasets/task/multimodal-emotion-recognition"},{"name":"Dialogue Act Classification","url":"/task/dialogue-act-classification","datasets_with_task":"/datasets/task/dialogue-act-classification"},{"name":"Open-Domain Dialog","url":"/task/open-domain-dialog","datasets_with_task":"/datasets/task/open-domain-dialog"},{"name":"Personality Trait Recognition","url":"/task/personality-trait-recognition","datasets_with_task":"/datasets/task/personality-trait-recognition"},{"name":"Dialog Act Classification","url":"/task/dialog-act-classification","datasets_with_task":"/datasets/task/dialog-act-classification"},{"name":"Conversational Response Generation","url":"/task/conversational-response-generation","datasets_with_task":"/datasets/task/conversational-response-generation"},{"name":"Personality Recognition in Conversation","url":"/task/personality-recognition-in-conversation","datasets_with_task":"/datasets/task/personality-recognition-in-conversation"},{"name":"Personalized and Emotional Conversation","url":"/task/personalized-and-emotional-conversation","datasets_with_task":"/datasets/task/personalized-and-emotional-conversation"},{"name":"Emotional Dialogue Acts","url":"/task/emotional-dialogue-acts","datasets_with_task":"/datasets/task/emotional-dialogue-acts"}],"languages":[{"name":"Chinese","url":"/datasets/language/chinese"}],"variants":["CPED"],"data_loaders":[{"repo":"https://github.com/scutcyr/CPED","url":"https://github.com/scutcyr/CPED","frameworks":["pytorch"]}],"num_papers_in_archive":15,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/emotion-recognition-in-conversation-on-cped","task":"Emotion Recognition in Conversation","dataset_variant":"CPED","rows":11,"metrics":["Accuracy of Sentiment","Macro-F1 of Sentiment"],"first_row_in_archive_order":{"model":"BERT+AVG+MLP","paper":"/paper/cped-a-large-scale-chinese-personalized-and-1","metrics":{"Accuracy of Sentiment":"51.50","Macro-F1 of Sentiment":"48.02"},"code_links":[{"title":"scutcyr/CPED","url":"https://github.com/scutcyr/CPED"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/personalized-and-emotional-conversation-on","task":"Personalized and Emotional Conversation","dataset_variant":"CPED","rows":9,"metrics":["PPL","BLEU","Distinct-1","Distinct-2","Greedy Embedding","Average Embedding","bertscore"],"first_row_in_archive_order":{"model":"GPT-{emo}","paper":"/paper/cped-a-large-scale-chinese-personalized-and-1","metrics":{"Average Embedding":"0.5588","BLEU":"0.1342","Distinct-1":"0.0614","Distinct-2":"0.3430","Greedy Embedding":"0.4996","PPL":"17.48","bertscore":"0.5709"},"code_links":[{"title":"scutcyr/CPED","url":"https://github.com/scutcyr/CPED"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/personality-recognition-in-conversation-on-1","task":"Personality Recognition in Conversation","dataset_variant":"CPED","rows":4,"metrics":["Accuracy (%)","Macro-F1","Accuracy of Neurotism","Accuracy of Extraversion","Accuracy of Openness","Accuracy of Agreeableness","Accuracy of Conscientiousness"],"first_row_in_archive_order":{"model":"BERT$_{ssenet}^{c}$","paper":"/paper/cped-a-large-scale-chinese-personalized-and-1","metrics":{"Accuracy (%)":"67.25","Accuracy of Agreeableness":"85.89","Accuracy of Conscientiousness":"63.48","Accuracy of Extraversion":"78.21","Accuracy of Neurotism":"53.27","Accuracy of Openness":"55.42","Macro-F1":"74.08"},"code_links":[{"title":"scutcyr/CPED","url":"https://github.com/scutcyr/CPED"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/cped-a-large-scale-chinese-personalized-and-1","title":"CPED: A Large-Scale Chinese Personalized and Emotional Dialogue Dataset for Conversational AI","date":"2022-05-29","rows_on_this_dataset":14,"code_links":1,"syntology":null},{"paper":"/paper/emoberta-speaker-aware-emotion-recognition-in","title":"EmoBERTa: Speaker-Aware Emotion Recognition in Conversation with RoBERTa","date":"2021-08-26","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/dialogxl-all-in-one-xlnet-for-multi-party","title":"DialogXL: All-in-One XLNet for Multi-Party Conversation Emotion Recognition","date":"2020-12-16","rows_on_this_dataset":1,"code_links":4,"syntology":null},{"paper":"/paper/dialoguegcn-a-graph-convolutional-neural","title":"DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation","date":"2019-08-30","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":4,"samples_ran":0,"samples_unverified":4,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/dialoguernn-an-attentive-rnn-for-emotion","title":"DialogueRNN: An Attentive RNN for Emotion Detection in Conversations","date":"2018-11-01","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/bert-pre-training-of-deep-bidirectional","title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","date":"2018-10-11","rows_on_this_dataset":1,"code_links":534,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":659,"samples_ran":208,"samples_unverified":451,"pointer_only_for_licence":149,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/context-dependent-sentiment-analysis-in-user","title":"Context-Dependent Sentiment Analysis in User-Generated Videos","date":"2017-07-01","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/bag-of-tricks-for-efficient-text","title":"Bag of Tricks for Efficient Text Classification","date":"2016-07-06","rows_on_this_dataset":1,"code_links":65,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":9,"samples_ran":2,"samples_unverified":7,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/recurrent-neural-network-for-text","title":"Recurrent Neural Network for Text Classification with Multi-Task Learning","date":"2016-05-17","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/recurrent-convolutional-neural-networks-for-2","title":"Recurrent Convolutional Neural Networks for Text Classification","date":"2015-01-01","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/convolutional-neural-networks-for-sentence","title":"Convolutional Neural Networks for Sentence Classification","date":"2014-08-25","rows_on_this_dataset":1,"code_links":118,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":77,"samples_ran":19,"samples_unverified":58,"pointer_only_for_licence":15,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":4,"samples_harvested":749,"samples_ran":229,"samples_unverified":520,"pointer_only_for_licence":166,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}