Datasets › Summarize from Feedback

Summarize from Feedback

Introduced by Nisan Stiennon et al. in Learning to summarize from human feedback2 Sep 2020 archive 2025-07-28

In the Learning to Summarize from Human Feedback paper, a reward model was trained from human feedback. The reward model was then used to train a summarization model to align with human preferences. This is the dataset of human feedback that was released for reward modelling. There are two parts of this dataset: comparisons and axis. In the comparisons part, human annotators were asked to choose the best out of two summaries. In the axis part, human annotators gave scores on a likert scale for the quality of a summary. The comparisons part only has a train and validation split, and the axis part only has a test and validation split.

Li et al. propose a variant with a subset of workers who annotate the data (details in Appendix C.1)

  1. Reddit TL;DR (Seen) uses the top 10 workers from the original dataset.
  2. Reddit TL;DR (Unseen) uses unseen workers in the validation set.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 5 papers for it but never published that list.

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Summarize from Feedback
  • Reddit TLDR (Seen)
  • Reddit TLDR (Unseen)

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections