Browse State-of-the-Art › Extreme Summarization
Extreme Summarization
13 papers with code · 4 benchmarks · 7 datasets archive 2025-07-28
Image credit: TLDR: Extreme Summarization of Scientific Documents
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
4 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CiteSum (9 rows) | EXT-ORACLE | CiteSum: Citation Text-guided Scientific Extreme Summarization and... | code | — | Compare |
| GEM-XSum (6 rows) | PEGASUS | The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics | — | — | Compare |
| TLDR9+ (4 rows) | ORACLE-EXT | TLDR9+: A Large Scale Resource for Extreme Summarization of Social... | code | — | Compare |
| XSum (1 row) | PEGASUS | The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
13 shown of 13 papers with code (22 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
28 May 2021 5 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedMost widely-used pre-trained language models operate on sequences of tokens corresponding to word or subword units.
-
9 Feb 2016 5 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedAttention mechanisms in neural networks have proved useful for problems in which the input and output do not have fixed dimension.
-
30 Apr 2020 4 repositories listed Syntology ran 12 of 19 samples · 7 unverified · 4 pointer-only (licence)We introduce TLDR generation, a new form of extreme summarization, for scientific papers.
-
7 Sep 2021 3 repositories listedWe present IndicBART, a multilingual, sequence-to-sequence pre-trained model focusing on 11 Indic languages and English.
-
27 Aug 2018 3 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedWe introduce extreme summarization, a new single-document summarization task which does not favor extractive strategies and calls for an abstractive modeling approach.
-
23 Dec 2024 1 repository listed Syntology ran 0 of 15 samples · 15 unverifiedFurthermore, when the size constraint of PKG is extremely small, the existing methods cannot distinguish which facts are more of immediate interest and guarantee the utility of the summarized PKG.
-
8 Mar 2024 1 repository listedKeywords, that is, content-relevant words in summaries play an important role in efficient information conveyance, making it critical to assess if system-generated summaries contain such informative words during…
-
27 Sep 2022 1 repository listedIn this paper, we introduce WikiDes, a novel dataset to generate short descriptions of Wikipedia articles for the problem of text summarization.
-
30 May 2022 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)The number of scientific publications nowadays is rapidly increasing, causing information overload for researchers and making it hard for scholars to keep up to date with current trends and lines of work.
-
12 May 2022 1 repository listedScientific extreme summarization (TLDR) aims to form ultra-short summaries of scientific papers.
-
4 Oct 2021 1 repository listedRecent models in developing summarization systems consist of millions of parameters and the model performance is highly dependent on the abundance of training data.
-
27 Oct 2020 1 repository listedMulti-document summarization is a challenging task for which there exists little large-scale datasets.
-
19 Jul 2019 1 repository listedWe introduce 'extreme summarization', a new single-document summarization task which aims at creating a short, one-sentence news summary answering the question ``What is the article about?''.
Syntology lines on 6 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections