{"url":"/dataset/peersum","name":"PeerSum","full_name":null,"description_markdown":"**PeerSum** is a new MDS dataset using peer reviews of scientific publications. The dataset differs from the existing MDS datasets in that summaries (i.e., the meta-reviews) are highly abstractive and they are real summaries of the source documents. \r\n\r\nIn PeerSum, we have reviews (with scores), comments and responses as the source documents and the meta-review (with an acceptance outcome) as the ground truth summary. Each sample of this dataset contains a summary, corresponding source documents and also other complementary information (e.g., review scores) for one paper. The second version of PeerSum (peersum_v2) has 16,308 samples, while there are 10,862 samples in the first version.\r\n\r\nThe dataset is stored in the json format. For each sample, details are based on following keys with explanation:\r\n\r\n* paper_id: unique id for each sample\r\n* title: the title of the corresponding paper\r\n* abstract: paper abstract\r\n* score: final score of this paper (if there is not a final, it will be an average of review scores)\r\n* acceptance: acceptance of the paper (e.g., accept, reject or spotlight)\r\n* meta_review: meta-review of the paper and this is treated as the summary\r\n* reviews: [review_id, writer, content (rating, confidence, comment), replyto]   review_id and replyto are for the conversation structure\r\n* label: train, val, test (8/1/1)\r\n\r\nFor each review (i.e., official review, public comment, or author/reviewer response):\r\n* review_id: unique id of each review\r\n* writer: official_reviewer, public, author\r\n* content: (rating, confidence, comment)\r\n* replyto: connect to a review (review_id and replyto are for the conversation structure)","description_withheld":null,"homepage":"https://github.com/oaimli/PeerSum","introduced_date":"2022-03-03","introduced_date_note":null,"introduced_by":{"paper":"/paper/peersum-a-peer-review-dataset-for-abstractive-1","title":"PeerSum: A Peer Review Dataset for Abstractive Multi-document Summarization","first_author":"Miao Li","url":null},"license":null,"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Abstractive Text Summarization","url":"/task/abstractive-text-summarization","datasets_with_task":"/datasets/task/abstractive-text-summarization"}],"languages":[],"variants":["PeerSum"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}