Datasets › Multi-News
Multi-News
Multi-News, consists of news articles and human-written summaries of these articles from the site newser.com. Each summary is professionally written by editors and includes links to the original articles cited.
Source: Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model Image Source: https://arxiv.org/pdf/1906.01749.pdf
Benchmarks archive 2025-07-28
All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Multi-Document Summarization | Multi-News | PRIMER ROUGE-1 49.9 | PRIMERA: Pyramid-based Masked Sentence Pre-training for... | allenai/primer +2 | 6 | Compare |
| Cross-Document Language Modeling | MultiNews val | CD-LM Perplexity 1.69 | CDLM: Cross-Document Language Modeling | aviclu/cdlm +1 | 3 | Compare |
| Cross-Document Language Modeling | MultiNews test | CD-LM Perplexity 1.76 | CDLM: Cross-Document Language Modeling | aviclu/cdlm +1 | 3 | Compare |
| Information Threading | Multi-News | SeqINT NMI 0.8008 | Identifying chronological and coherent information... | hitt08/HINT | 1 | Compare |
| Summarization | multi_news | no rows | — | — | 0 | Compare |
Papers archive 2025-07-28
7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 122. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Identifying chronological and coherent information threads using 5W1H questions and temporal relationships | 1 | 1 | 18 Jan 2023 | not harvested |
| LongT5: Efficient Text-To-Text Transformer for Long Sequences | 4 | 1 | 15 Dec 2021 | ran 1 of 1 samples (0 unverified; 1 pointer-only for licence) |
| PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document Summarization | 3 | 1 | 16 Oct 2021 | ran 4 of 7 samples (3 unverified) |
| Multi-Document Summarization withDeterminantal Point Process Attention | 0 | 1 | 13 Jul 2021 | not harvested |
| CDLM: Cross-Document Language Modeling | 2 | 6 | 2 Jan 2021 | not harvested |
| Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model | 1 | 1 | 4 Jun 2019 | not harvested |
| Bottom-Up Abstractive Summarization | 5 | 2 | 31 Aug 2018 | ran 5 of 6 samples (1 unverified) |
Dataset loaders archive 2025-07-28
11 loaders as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- MultiNews val
- MultiNews test
- Multi-News
- multi_news
4 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections