{"url":"/dataset/oposum","name":"OpoSum","full_name":null,"description_markdown":"OPOSUM is a dataset for the training and evaluation of Opinion Summarization models which contains Amazon reviews from six product domains: Laptop Bags, Bluetooth Headsets, Boots, Keyboards, Televisions, and Vacuums.\r\nThe six training collections were created by downsampling from the Amazon Product Dataset introduced in McAuley et al. (2015) and contain reviews and their respective ratings. \r\n\r\nA subset of the dataset has been manually annotated, specifically, for each domain, 10 different products were uniformly sampled (across ratings) with 10 reviews each, amounting to a total of 600 reviews, to be used only for development (300) and testing (300).\r\n\r\nSource: [Summarizing Opinions: Aspect Extraction Meets Sentiment Prediction and They Are Both Weakly Supervised](https://arxiv.org/abs/1808.08858)","description_withheld":null,"homepage":"https://github.com/stangelid/oposum","introduced_date":"2018-08-27","introduced_date_note":null,"introduced_by":{"paper":"/paper/summarizing-opinions-aspect-extraction-meets","title":"Summarizing Opinions: Aspect Extraction Meets Sentiment Prediction and They Are Both Weakly Supervised","first_author":"Stefanos Angelidis","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Topic Models","url":"/task/topic-models","datasets_with_task":"/datasets/task/topic-models"},{"name":"Document Summarization","url":"/task/document-summarization","datasets_with_task":"/datasets/task/document-summarization"},{"name":"Multi-Document Summarization","url":"/task/multi-document-summarization","datasets_with_task":"/datasets/task/multi-document-summarization"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["OpoSum"],"data_loaders":[{"repo":"https://github.com/stangelid/oposum","url":"https://github.com/stangelid/oposum","frameworks":["pytorch"]}],"num_papers_in_archive":7,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}