Browse State-of-the-Art › Extractive Text Summarization
Extractive Text Summarization
34 papers with code · 5 benchmarks · 5 datasets archive 2025-07-28
Given a document, selecting a subset of the words or sentences which best represents a summary of the document.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CNN / Daily Mail (15 rows) | HAHSum | Neural Extractive Summarization with Hierarchical Attentive... | — | — | Compare |
| DebateSum (3 rows) | Longformer-Base | DebateSum: A large-scale argument mining and summarization dataset | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| GovReport (2 rows) | MemSum (extractive) | MemSum: Extractive Summarization of Long Documents Using... | code | — | Compare |
| DUC 2004 Task 1 (1 row) | Abs | A Neural Attention Model for Abstractive Sentence Summarization | code | — | Compare |
| DUC 2004 (1 row) | Pre-training-meets-Clustering-A-Hybrid-Extractive-Multi-Document-Summarization-Model | Pre-training Meets Clustering: A Hybrid Extractive Multi-document... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
5 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 34 papers with code (95 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
14 Apr 2017 39 repositories listed Syntology ran 30 of 64 samples · 34 unverified · 44 pointer-only (licence)Neural sequence-to-sequence models have provided a viable new approach for abstractive text summarization (meaning they are not restricted to simply selecting and rearranging passages from the original text).
-
22 Aug 2019 19 repositories listed Syntology ran 7 of 21 samples · 14 unverifiedFor abstractive summarization, we propose a new fine-tuning schedule which adopts different optimizers for the encoder and the decoder as a means of alleviating the mismatch between the two (the former is pretrained…
-
4 Dec 2018 14 repositories listed Syntology ran 7 of 8 samples · 1 unverified · 1 pointer-only (licence)Dot-product attention has wide applications in computer vision and natural language processing.
-
25 Mar 2019 12 repositories listed Syntology ran 2 of 5 samples · 3 unverifiedBERT, a pre-trained Transformer model, has achieved ground-breaking performance on multiple NLP tasks.
-
7 Jun 2019 8 repositories listedThis paper reports on the project called Lecture Summarization Service, a python based RESTful service that utilizes the BERT model for text embeddings and KMeans clustering to identify sentences closes to the centroid…
-
2 Sep 2015 4 repositories listedSummarization based on text extraction is inherently limited, but generation-style abstractive methods have proven challenging to build.
-
14 Nov 2020 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Finally, we present a search engine for this dataset which is utilized extensively by members of the National Speech and Debate Association today.
-
13 Apr 2020 3 repositories listedRedundancy-aware extractive summarization systems score the redundancy of the sentences to be included in a summary either jointly with their salience information or separately as an additional sentence scoring step.
-
13 Mar 2023 2 repositories listedThe goal of multimodal summarization is to extract the most important information from different modalities to form output summaries.
-
7 Dec 2020 2 repositories listedCompetitive Debate's increasingly technical nature has left competitors looking for tools to accelerate evidence production.
-
27 Apr 2020 2 repositories listedMost general-purpose extractive summarization models are trained on news articles, which are short and present all important information upfront.
-
19 Apr 2020 2 repositories listedThis paper creates a paradigm shift with regard to the way we build neural extractive summarization systems.
-
8 Jul 2019 2 repositories listedThe recent years have seen remarkable success in the use of deep neural networks on text summarization.
-
20 Feb 2018 2 repositories listedDetecting novelty of an entire document is an Artificial Intelligence (AI) frontier problem that has widespread NLP applications, such as extractive document summarization, tracking development of news events,…
-
1 Apr 2017 2 repositories listedThe textual similarity is a crucial aspect for many extractive text summarization methods.
-
26 Nov 2024 1 repository listedHere, we propose a novel Word pair-based Gaussian Sentence Similarity (WGSS) algorithm for calculating the semantic relation between two sentences.
-
25 May 2023 1 repository listedOutcomes validate that our proposed model shows greatly enhanced performance as compared to the existent unsupervised state-of-the-art approaches.
-
1 Oct 2022 1 repository listedIn this paper, we develop a Graph-Based Unsupervised Summarization(GUSUM) method for extractive text summarization based on the principle of including the most important sentences while excluding sentences with similar…
-
1 Oct 2022 1 repository listedIn this paper, we describe the first IDN dataset (IDN-Sum) designed specifically for training and testing IDN text summarization algorithms.
-
19 Jul 2021 1 repository listedWe introduce MemSum (Multi-step Episodic Markov decision process extractive SUMmarizer), a reinforcement-learning-based extractive summarizer enriched at each step with information on the current extraction history.
-
16 Oct 2020 1 repository listed Syntology ran 0 of 3 samples · 3 unverified · 3 pointer-only (licence)We also find in experiments that our model is less dependent on sentence positions.
-
26 Apr 2020 1 repository listedAn intuitive way is to put them in the graph-based neural network, which has a more complex structure for capturing inter-sentence relationships.
-
4 Dec 2019 1 repository listedMulti-document summarization is more challenging than single-document summarization since it has to solve the problem of overlapping information among sentences from different documents.
-
1 Nov 2019 1 repository listedIn this work, we re-examine the problem of extractive text summarization for long documents.
-
30 Oct 2019 1 repository listed Syntology ran 1 of 11 samples · 10 unverifiedRecently BERT has been adopted for document encoding in state-of-the-art text summarization models.
-
16 Jul 2019 1 repository listedOur method creates an extractive summary by selecting the sentences with the closest embeddings to the document embedding.
-
3 Feb 2019 1 repository listedIn this work, we present a neural model for single-document summarization based on joint extraction and syntactic compression.
-
6 Nov 2018 1 repository listedWe propose DeepChannel, a robust, data-efficient, and interpretable neural model for extractive document summarization.
-
27 Sep 2018 1 repository listedIn this paper, we introduce Iterative Text Summarization (ITS), an iteration-based model for supervised extractive text summarization, inspired by the observation that it is often necessary for a human to read an…
-
25 Sep 2018 1 repository listedIn this work, we propose a novel method for training neural networks to perform single-document extractive summarization without heuristically-generated extractive labels.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections