{"url":"/dataset/summarize-from-feedback","name":"Summarize from Feedback","full_name":null,"description_markdown":"In the Learning to Summarize from Human Feedback paper, a reward model was trained from human feedback. The reward model was then used to train a summarization model to align with human preferences. This is the dataset of human feedback that was released for reward modelling. There are two parts of this dataset: comparisons and axis. In the comparisons part, human annotators were asked to choose the best out of two summaries. In the axis part, human annotators gave scores on a likert scale for the quality of a summary. The comparisons part only has a train and validation split, and the axis part only has a test and validation split.\r\n\r\nLi et al. propose a variant with a subset of workers who annotate the data (details in [Appendix C.1](https://arxiv.org/pdf/2402.05133)) \r\n\r\n1. Reddit TL;DR (Seen) uses the top 10 workers from the original dataset.\r\n2. Reddit TL;DR (Unseen) uses unseen workers in the validation set.","description_withheld":null,"homepage":"https://github.com/openai/summarize-from-feedback","introduced_date":"2020-09-02","introduced_date_note":null,"introduced_by":{"paper":"/paper/learning-to-summarize-from-human-feedback","title":"Learning to summarize from human feedback","first_author":"Nisan Stiennon","url":null},"license":null,"modalities":[],"tasks":[{"name":"Preference Mapping","url":"/task/preference-mapping","datasets_with_task":"/datasets/task/preference-mapping"}],"languages":[],"variants":["Summarize from Feedback","Reddit TLDR (Seen)","Reddit TLDR (Unseen)"],"data_loaders":[{"repo":"https://github.com/openai/summarize-from-feedback","url":"https://github.com/openai/summarize-from-feedback","frameworks":[]}],"num_papers_in_archive":5,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}