{"url":"/dataset/dialogsum","name":"DialogSum","full_name":null,"description_markdown":"DialogSum is a large-scale dialogue summarization dataset, consisting of 13,460 dialogues with corresponding manually labeled summaries and topics.\r\n\r\nThis work is accepted by ACL findings 2021. You may find the paper here: <https://arxiv.org/pdf/2105.06762.pdf>.\r\n\r\nIf you want to use our dataset, please cite our paper. \r\n\r\n#### Dialogue Data\r\n\r\nWe collect dialogue data for DialogSum from three public dialogue corpora, namely Dailydialog (Li et al., 2017), DREAM (Sun et al., 2019) and MuTual (Cui et al., 2019), as well as an English speaking practice website. \r\nThese datasets contain face-to-face spoken dialogues that cover a wide range of daily-life topics, including schooling, work, medication, shopping, leisure, travel.\r\nMost conversations take place between friends, colleagues, and between service providers and customers.\r\n\r\nCompared with previous datasets, dialogues from DialogSum have distinct characteristics: \r\n* Under rich real-life scenarios, including more diverse task-oriented scenarios;\r\n* Have clear communication patterns and intents, which is valuable to serve as summarization sources;\r\n* Have a reasonable length, which comforts the purpose of automatic summarization.\r\n\r\n#### Summaries\r\n\r\nWe ask annotators to summarize each dialogue based on the following criteria:\r\n* Convey the most salient information;\r\n* Be brief;\r\n* Preserve important named entities within the conversation;\r\n* Be written from an observer perspective;\r\n* Be written in formal language.\r\n\r\n#### Topics\r\nIn addition to summaries, we also ask annotators to write a short topic for each dialogue, which can be potentially useful for future work, e.g. generating summaries by leveraging topic information.\r\n\r\nImage source: [https://arxiv.org/pdf/2105.06762.pdf](https://arxiv.org/pdf/2105.06762.pdf)","description_withheld":null,"homepage":"https://github.com/cylnlp/DialogSum","introduced_date":"2021-05-14","introduced_date_note":null,"introduced_by":{"paper":"/paper/dialsumm-a-real-life-scenario-dialogue","title":"DialogSum: A Real-Life Scenario Dialogue Summarization Dataset","first_author":"Yulong Chen","url":null},"license":{"name":"MIT","url":"https://github.com/cylnlp/DialogSum/blob/main/LICENSE"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Text Summarization","url":"/task/text-summarization","datasets_with_task":"/datasets/task/text-summarization"},{"name":"Abstractive Text Summarization","url":"/task/abstractive-text-summarization","datasets_with_task":"/datasets/task/abstractive-text-summarization"},{"name":"Dialogue Generation","url":"/task/dialogue-generation","datasets_with_task":"/datasets/task/dialogue-generation"}],"languages":[{"name":"Chinese","url":"/datasets/language/chinese"}],"variants":["DialogSum"],"data_loaders":[],"num_papers_in_archive":62,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/text-summarization-on-dialogsum","task":"Text Summarization","dataset_variant":"DialogSum","rows":4,"metrics":["Rouge1","Rouge2","RougeL","BertScore"],"first_row_in_archive_order":{"model":"InstructDS","paper":"/paper/instructive-dialogue-summarization-with-query","metrics":{"Rouge1":"47.8","Rouge2":"22.2","RougeL":"39.4"},"code_links":[{"title":"BinWang28/InstructDS","url":"https://github.com/BinWang28/InstructDS"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/abstractive-text-summarization-on-dialogsum","task":"Abstractive Text Summarization","dataset_variant":"DialogSum","rows":0,"metrics":["Test ROGUE-1","Test ROGUE-2","Test ROGUE-L","Test ROGUE-Lsum","Validation ROGUE-1","Validation ROGUE-2","Validation ROGUE-L","Validation ROGUE-Lsum"],"first_row_in_archive_order":null,"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/omnivec2-a-novel-transformer-based-network","title":"OmniVec2 - A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning","date":"2024-01-01","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/omnivec-learning-robust-representations-with","title":"OmniVec: Learning robust representations with cross modal sharing","date":"2023-11-07","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/instructive-dialogue-summarization-with-query","title":"Instructive Dialogue Summarization with Query Aggregations","date":"2023-10-17","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/mind-the-gap-injecting-commonsense-knowledge","title":"Mind the Gap! Injecting Commonsense Knowledge for Abstractive Dialogue Summarization","date":"2022-09-02","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}