{"url":"/dataset/subsume","name":"SubSumE","full_name":null,"description_markdown":"# SubSumE Dataset\r\n\r\nThis repository contains the SubSumE dataset for subjective document summarization. See [the paper](https://aclanthology.org/2021.newsum-1.14/) and the [talk](https://www.youtube.com/watch?v=0vyUQArRrvY) for details on dataset creation. Also check out our work [SuDocu](http://sudocu.cs.umass.edu/) on example-based document summarization.\r\n\r\n\r\n## Dataset Files\r\nDownload the dataset from [here](https://drive.google.com/file/d/1tEDDHzZM_idnv-_PfRE5BmJU5E8yKLRH/view).\r\n\r\n\r\nThe dataset contains :\r\n\r\n* Simplified text from 48 Wikipedia pages of the states in the US. Additionally, all the sentences in these documents\r\nare put together in a single file `processed_state_sentences.csv` and are assigned a unique sentence id that \r\nis used in summary json files. \r\n* Intent-based summaries created by human annotators.\r\n\r\nEach datapoint file in the directory `user_summary_jsons` contains a json containing summaries of Wikipedia pages\r\nof eight states with following keys:\r\n\r\n* **intent** : Summarization intent provided to human annotators for generating the summary\r\n* **summaries**: List of summary jsons for eight states assigned to the annotator. Each json in the list contains following keys:\r\n    * **state_name**: Name of the state\r\n    * **sentence_ids**: Global ids of sentences (wrt `processed_state_sentences.csv`) present in the summary\r\n    * **sentences**: List of sentences present in the summary\r\n    * **use_keywords**: Keywords used by the annotator to search the document when creating summaries","description_withheld":null,"homepage":"https://github.com/afariha/SubSumE","introduced_date":"2021-11-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/subsume-a-dataset-for-subjective-summary","title":"SUBSUME: A Dataset for Subjective Summary Extraction from Wikipedia Documents","first_author":"Nishant Yadav","url":null},"license":{"name":"CC-BY-4.0 License","url":"https://github.com/afariha/SubSumE/blob/main/LICENSE"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Text Summarization","url":"/task/text-summarization","datasets_with_task":"/datasets/task/text-summarization"},{"name":"Document Summarization","url":"/task/document-summarization","datasets_with_task":"/datasets/task/document-summarization"},{"name":"Extractive Text Summarization","url":"/task/extractive-document-summarization","datasets_with_task":"/datasets/task/extractive-document-summarization"},{"name":"Query-Based Extractive Summarization","url":"/task/query-based-extractive-summarization","datasets_with_task":"/datasets/task/query-based-extractive-summarization"},{"name":"Extractive Document Summarization","url":"/task/extractive-document-summarization-1","datasets_with_task":"/datasets/task/extractive-document-summarization-1"},{"name":"Extractive Summarization","url":"/task/extractive-summarization","datasets_with_task":"/datasets/task/extractive-summarization"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["SubSumE"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}