Papers › Discourse-Aware Unsupervised Summarization of Long Scientific Documents

Discourse-Aware Unsupervised Summarization of Long Scientific Documents

1 May 2020arXiv:2005.00513archive 2025-07-28

Yue Dong, Andrei Mircea, Jackie C. K. Cheung

We propose an unsupervised graph-based ranking model for extractive summarization of long scientific documents. Our method assumes a two-level hierarchical graph representation of the source document, and exploits asymmetrical positional cues to determine sentence importance. Results on the PubMed and arXiv datasets show that our approach outperforms strong unsupervised baselines by wide margins in automatic metrics and human evaluation. In addition, it achieves performance comparable to many state-of-the-art supervised approaches which are trained on hundreds of thousands of examples. These results suggest that patterns in the discourse structure are a strong signal for determining importance in scientific articles.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

mirandrom/HipoRank mentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ArticlesExtractive SummarizationSentenceUnsupervised Extractive Summarization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Unsupervised Extractive Summarization Pubmed HipoRank ROUGE-1 43.58 #1 of 9 Archive leaderboard report
Unsupervised Extractive Summarization Pubmed HipoRank ROUGE-2 17.00 #1 of 9 Archive leaderboard report
Unsupervised Extractive Summarization Pubmed HipoRank ROUGE-L 39.31 #1 of 9 Archive leaderboard report
Unsupervised Extractive Summarization Pubmed PacSum ROUGE-1 39.79 #3 of 9 Archive leaderboard report
Unsupervised Extractive Summarization Pubmed PacSum ROUGE-2 14.00 #3 of 9 Archive leaderboard report
Unsupervised Extractive Summarization Pubmed PacSum ROUGE-L 36.09 #3 of 9 Archive leaderboard report
Unsupervised Extractive Summarization arXiv Summarization Dataset HipoRank ROUGE-1 39.34 #2 of 10 Archive leaderboard report
Unsupervised Extractive Summarization arXiv Summarization Dataset HipoRank ROUGE-2 12.56 #2 of 10 Archive leaderboard report
Unsupervised Extractive Summarization arXiv Summarization Dataset HipoRank ROUGE-L 34.89 #2 of 10 Archive leaderboard report
Unsupervised Extractive Summarization arXiv Summarization Dataset PacSum ROUGE-1 38.57 #4 of 10 Archive leaderboard report
Unsupervised Extractive Summarization arXiv Summarization Dataset PacSum ROUGE-2 10.93 #4 of 10 Archive leaderboard report
Unsupervised Extractive Summarization arXiv Summarization Dataset PacSum ROUGE-L 34.33 #4 of 10 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBERTDense ConnectionsDropoutLayer NormalizationLinear LayerLinear Warmup With Linear DecayMulti-Head AttentionResidual ConnectionRoBERTaSoftmaxWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections