Browse State-of-the-Art › Text Summarization

Text Summarization

440 papers with code · 37 benchmarks · 98 datasets archive 2025-07-28

Knowledge BaseNatural Language Processing

Text Summarization is a natural language processing (NLP) task that involves condensing a lengthy text document into a shorter, more compact version while still retaining the most important information and meaning. The goal is to produce a summary that accurately represents the content of the original text in a concise form.

There are different approaches to text summarization, including extractive methods that identify and extract important sentences or phrases from the text, and abstractive methods that generate new text based on the content of the original text.

Description from the archive archive 2025-07-28.

Benchmarks archive 2025-07-28

37 leaderboard tables shown for this task, 37 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 37 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
GigaWord (41 rows) OpenAI/o3-mini — — — Compare
Pubmed (29 rows) Top Down Transformer (AdaPool) (464M) Long Document Summarization with Top-down and Bottom-up Inference code — Compare
Arxiv HEP-TH citation graph (28 rows) Top Down Transformer (AdaPool) (464M) Long Document Summarization with Top-down and Bottom-up Inference code — Compare
MTEB (26 rows) MPNet-multilingual MTEB: Massive Text Embedding Benchmark code Syntology ran 3 of 13 samples · 10 unverified Compare
X-Sum (18 rows) Selfmem Lift Yourself Up: Retrieval-augmented Text Generation with Self Memory code Syntology ran 1 of 1 samples · 0 unverified Compare
CNN / Daily Mail (Anonymized) (13 rows) HSSAS A Hierarchical Structured Self-Attentive Model for Extractive... — — Compare
DUC 2004 Task 1 (13 rows) Transformer+WDrop Rethinking Perturbations in Encoder-Decoders for Fast Training code Syntology ran 1 of 1 samples · 0 unverified Compare
SAMSum (12 rows) OmniVec2 OmniVec2 - A Novel Transformer based Network for Large Scale... — — Compare
Reddit TIFU (5 rows) PEGASUS 2B + SLiC Calibrating Sequence likelihood Improves Conditional Language Generation — — Compare
arXiv Summarization Dataset (4 rows) PRIMER PRIMERA: Pyramid-based Masked Sentence Pre-training for... code Syntology ran 4 of 7 samples · 3 unverified Compare
DialogSum (4 rows) InstructDS Instructive Dialogue Summarization with Query Aggregations code — Compare
Klexikon (4 rows) Luhn's algorithm (25 sentences) Klexikon: A German Dataset for Joint Summarization and Simplification code — Compare
BookSum (3 rows) Echoes-Extractive-Abstractive Echoes from Alexandria: A Large Resource for Multilingual Book... code — Compare
GigaWord-10k (3 rows) ERNIE-GENLARGE (large-scale text corpora) ERNIE-GEN: An Enhanced Multi-Flow Pre-training and Fine-tuning... code — Compare
WikiHow (3 rows) BertSum Abstractive Summarization of Spoken andWritten Instructions with BERT code — Compare
BigPatent (2 rows) LongT5 LongT5: Efficient Text-To-Text Transformer for Long Sequences code Syntology ran 1 of 1 samples · 0 unverified Compare
GovReport (2 rows) FactorSum Factorizing Content and Budget Decisions in Abstractive... code — Compare
How2 (2 rows) Ground-truth transcript + Action with Hierarchical Attn Multimodal Abstractive Summarization for How2 Videos — — Compare
MeetingBank (2 rows) CriSPO 3-shot CriSPO: Multi-Aspect Critique-Suggestion-guided Automatic Prompt... code — Compare
OrangeSum (2 rows) mBARThez (OrangeSum abstract) BARThez: a Skilled Pretrained French Sequence-to-Sequence Model code — Compare
ACI-Bench (1 row) CriSPO 3-shot CriSPO: Multi-Aspect Critique-Suggestion-guided Automatic Prompt... code — Compare
AMI (1 row) HAT-CNNDM Hierarchical Learning for Generation with Long Source Sequences — — Compare
arXiv (1 row) BigBird-Pegasus Big Bird: Transformers for Longer Sequences code Syntology ran 10 of 15 samples · 5 unverified Compare
BBC XSum (1 row) MatchSum Extractive Summarization as Text Matching code — Compare
BillSum (1 row) Longformer Encoder Decoder BillSum: A Corpus for Automatic Summarization of US Legislation code — Compare
CL-SciSumm (1 row) GCN Hybrid ScisummNet: A Large Annotated Corpus and Content-Impact Models for... code Syntology ran 1 of 1 samples · 0 unverified Compare
CORD-19 (1 row) GenCompareSum GenCompareSum: a hybrid unsupervised summarization method using salience code — Compare
EurekaAlert (1 row) SATS SATS: simplification aware text summarization of scientific documents — — Compare
Gazeta (1 row) Finetuned mBART Dataset for Automatic Summarization of Russian News code — Compare
LCSTS (1 row) LSTM-seq2seq LCSTS: A Large Scale Chinese Short Text Summarization Dataset code Syntology ran 0 of 3 samples · 3 unverified Compare
MediaSum (1 row) SRformer-BART Segmented Recurrent Transformer: An Efficient Sequence-to-Sequence Model code — Compare
MentSum (1 row) BART MentSum: A Resource for Exploring Summarization of Mental Health... — — Compare
MeQSum (1 row) BiomedGPT BiomedGPT: A Generalist Vision-Language Foundation Model for... code Syntology ran 1 of 8 samples · 7 unverified Compare
QMSum (1 row) BART-LS Adapting Pretrained Text-to-Text Models for Long Text Sequences code — Compare
S2ORC (1 row) GenCompareSum GenCompareSum: a hybrid unsupervised summarization method using salience code — Compare
Webis-Snippet-20 Corpus (1 row) Anchor-context + Query biased Abstractive Snippet Generation code — Compare
XSum (1 row) SRformer-BART Segmented Recurrent Transformer: An Efficient Sequence-to-Sequence Model code — Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

98 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 98 until expanded.

Subtasks archive 2025-07-28

12 subtasks in the archive's task tree.

Most implemented papers archive 2025-07-28

30 shown of 440 papers with code (1,340 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

Syntology lines on 19 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections