{"url":"/dataset/discofuse","name":"DiscoFuse","full_name":"DiscoFuse","description_markdown":"DiscoFuse was created by applying a rule-based splitting method on two corpora -\r\nsports articles crawled from the Web, and Wikipedia. See the paper for a detailed\r\ndescription of the dataset generation process and evaluation.\r\n\r\nDiscoFuse has two parts with 44,177,443 and 16,642,323 examples sourced from Sports articles and Wikipedia, respectively.\r\n\r\nFor each part, a random split is provided to train (98% of the examples), development (1%) and test (1%) sets. In addition, as the original data distribution is highly skewed (see details in the paper), a balanced version for each part is also provided.\r\n\r\nSource: [Google Research](https://github.com/google-research-datasets/discofuse)","description_withheld":null,"homepage":"https://github.com/google-research-datasets/discofuse","introduced_date":"2019-02-27","introduced_date_note":null,"introduced_by":{"paper":"/paper/discofuse-a-large-scale-dataset-for-discourse","title":"DiscoFuse: A Large-Scale Dataset for Discourse-Based Sentence Fusion","first_author":"Mor Geva","url":null},"license":{"name":"CC BY-SA 3.0","url":"https://creativecommons.org/licenses/by-sa/3.0/"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Sentence Fusion","url":"/task/sentence-fusion","datasets_with_task":"/datasets/task/sentence-fusion"}],"languages":[],"variants":["DiscoFuse"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/google-research-datasets/discofuse","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/discofuse","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/google-research-datasets/discofuse","url":"https://github.com/google-research-datasets/discofuse","frameworks":[]}],"num_papers_in_archive":10,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}