Papers › Pre-training Meets Clustering: A Hybrid Extractive Multi-document Summarization Model

Pre-training Meets Clustering: A Hybrid Extractive Multi-document Summarization Model

25 May 2023International Conference on Hybrid Intelligent Systems 2023 5archive 2025-07-28

Akanksha Karotia, Seba Susan

In this era where a large amount of information has flooded the Internet, manual extraction and consumption of relevant information is very difficult and time-consuming. Therefore, an automated document summarization tool is necessary to excerpt important information from a set of documents that have similar or related subjects. Multi-document summarization allows retrieval of important and relevant content from multiple documents while minimizing redundancy. A multi-document text summarization system is developed in this study using an unsupervised extractive-based approach. The proposed model is a fusion of two learning paradigms: the T5 pre-trained transformer model and the K-Means clustering algorithm. We perform the experiments on the benchmark news article corpus Document Understanding Conference (DUC2004). The ROUGE evaluation metrics were used to estimate the performance of the proposed approach on the DUC2004. Outcomes validate that our proposed model shows greatly enhanced performance as compared to the existent unsupervised state-of-the-art approaches.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClusteringDocument SummarizationExtractive Text SummarizationMulti-Document SummarizationRetrievalText SummarizationUnsupervised Text Summarizationdocument understanding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Extractive Text Summarization DUC 2004 Pre-training-meets-Clustering-A-Hybrid-Extractive-Multi-Document-Summarization-Model Test ROGUE-1 34.013 #1 of 1 Archive leaderboard report
Extractive Text Summarization DUC 2004 Pre-training-meets-Clustering-A-Hybrid-Extractive-Multi-Document-Summarization-Model Test ROGUE-2 8.266 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdafactorAttentionAttention DropoutBPEDense ConnectionsDropoutGated Linear UnitInverse Square Root ScheduleLayer NormalizationLinear LayerMulti-Head AttentionResidual ConnectionSentencePieceSoftmaxT5k-Means Clustering

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections