Papers › Hierarchical Learning for Generation with Long Source Sequences

Hierarchical Learning for Generation with Long Source Sequences

15 Apr 2021arXiv:2104.07545archive 2025-07-28

Tobias Rohde, Xiaoxia Wu, Yinhan Liu

One of the challenges for current sequence to sequence (seq2seq) models is processing long sequences, such as those in summarization and document level machine translation tasks. These tasks require the model to reason at the token level as well as the sentence and paragraph level. We design and study a new Hierarchical Attention Transformer-based architecture (HAT) that outperforms standard Transformers on several sequence to sequence tasks. Furthermore, our model achieves state-of-the-art ROUGE scores on four summarization tasks, including PubMed, arXiv, CNN/DM, SAMSum, and AMI. Our model outperforms document-level machine translation baseline on the WMT20 English to German translation task. We investigate what the hierarchical layers learn by visualizing the hierarchical encoder-decoder attention. Finally, we study hierarchical learning on encoder-only pre-training and analyze its performance on classification tasks.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

DecoderDocument Level Machine TranslationDocument SummarizationDocument TranslationGeneral ClassificationMachine TranslationReading ComprehensionSentenceText SummarizationTranslation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Document Summarization CNN / Daily Mail HAT-BART ROUGE-1 44.48 #5 of 26 Archive leaderboard report
Document Summarization CNN / Daily Mail HAT-BART ROUGE-2 21.31 #5 of 26 Archive leaderboard report
Document Summarization CNN / Daily Mail HAT-BART ROUGE-L 41.52 #5 of 26 Archive leaderboard report
Reading Comprehension RACE HAT (Encoder) Accuracy 67.3 #10 of 24 Archive leaderboard report
Text Summarization AMI HAT-CNNDM ROUGE-1 52.27 #1 of 1 Archive leaderboard report
Text Summarization AMI HAT-CNNDM ROUGE-2 20.15 #1 of 1 Archive leaderboard report
Text Summarization AMI HAT-CNNDM ROUGE-L 50.57 #1 of 1 Archive leaderboard report
Text Summarization Arxiv HEP-TH citation graph HAT-BART ROUGE-1 46.74 #13 of 28 Archive leaderboard report
Text Summarization Arxiv HEP-TH citation graph HAT-BART ROUGE-2 19.19 #13 of 28 Archive leaderboard report
Text Summarization Arxiv HEP-TH citation graph HAT-BART ROUGE-L 42.2 #13 of 28 Archive leaderboard report
Text Summarization Pubmed HAT-BART ROUGE-1 48.25 #9 of 29 Archive leaderboard report
Text Summarization Pubmed HAT-BART ROUGE-2 21.35 #9 of 29 Archive leaderboard report
Text Summarization Pubmed HAT-BART ROUGE-L 36.69 #9 of 29 Archive leaderboard report
Text Summarization SAMSum HAT-CNNDM ROUGE-1 53.01 #7 of 12 Archive leaderboard report
Text Summarization SAMSum HAT-CNNDM ROUGE-2 28.27 #7 of 12 Archive leaderboard report
Text Summarization SAMSum HAT-CNNDM RL ROUGE-L 48.84 #11 of 12 Archive leaderboard report
Text Summarization X-Sum HAT-BART ROUGE-1 45.92 #6 of 18 Archive leaderboard report
Text Summarization X-Sum HAT-BART ROUGE-2 22.79 #6 of 18 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections