Methods › Natural Language Processing › Transformers › PEGASUS
PEGASUS
Introduced by Jingqing Zhang et al. in PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
PEGASUS proposes a transformer-based model for abstractive summarization. It uses a special self-supervised pre-training objective called gap-sentences generation (GSG) that's designed to perform well on summarization-related downstream tasks. As reported in the paper, "both GSG and MLM are applied simultaneously to this example as pre-training objectives. Originally there are three sentences. One sentence is masked with [MASK1] and used as target generation text (GSG). The other two sentences remain in the input, but some tokens are randomly masked by [MASK2]."
Papers archive 2025-07-28
30 shown of 53, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Pegasus: A Universal Framework for Scalable Deep Learning Inference on the Dataplane 6 Jun 2025 · 0 repositories · arXiv:2506.05779
-
QUAD-LLM-MLTC: Large Language Models Ensemble Learning for Healthcare Text Multi-Label Classification 20 Feb 2025 · 0 repositories · arXiv:2502.14189
-
Implementing Large Quantum Boltzmann Machines as Generative AI Models for Dataset Balancing 5 Feb 2025 · 0 repositories · arXiv:2502.03086
-
Extract-and-Abstract: Unifying Extractive and Abstractive Summarization within Single Encoder-Decoder Framework 18 Sep 2024 · 0 repositories · arXiv:2409.11827
-
GLIMMER: Incorporating Graph and Lexical Features in Unsupervised Multi-Document Summarization 19 Aug 2024 · 1 repository · arXiv:2408.10115
-
BioLay_AK_SS at BioLaySumm: Domain Adaptation by Two-Stage Fine-Tuning of Large Language Models used for Biomedical Lay Summary Generation 16 Aug 2024 · 0 repositories
-
Factual Dialogue Summarization via Learning from Large Language Models 20 Jun 2024 · 0 repositories · arXiv:2406.14709
-
Comparing Quantum Annealing and Spiking Neuromorphic Computing for Sampling Binary Sparse Coding QUBO Problems 30 May 2024 · 0 repositories · arXiv:2405.20525
-
Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT 7 May 2024 · 0 repositories · arXiv:2405.04053
-
MEDVOC: Vocabulary Adaptation for Fine-tuning Pre-trained Language Models on Medical Text Summarization 7 May 2024 · 1 repository · arXiv:2405.04163Syntology ran 8 of 9 samples · 1 unverified · 9 pointer-only (licence)
-
Analysis of Multidomain Abstractive Summarization Using Salience Allocation 19 Feb 2024 · 0 repositories · arXiv:2402.11955
-
PEGASUS: Personalized Generative 3D Avatars with Composable Attributes 16 Feb 2024 · 0 repositories · arXiv:2402.10636
-
Source Identification in Abstractive Summarization 7 Feb 2024 · 1 repository · arXiv:2402.04677
-
PEGASUS: Physically Enhanced Gaussian Splatting Simulation System for 6DoF Object Pose Dataset Generation 4 Jan 2024 · 1 repository · arXiv:2401.02281
-
Revisiting Zero-Shot Abstractive Summarization in the Era of Large Language Models from the Perspective of Position Bias 3 Jan 2024 · 1 repository · arXiv:2401.01989
-
Harnessing the Power of Prompt-based Techniques for Generating School-Level Questions using Large Language Models 2 Dec 2023 · 1 repository · arXiv:2312.01032
-
FaMeSumm: Investigating and Improving Faithfulness of Medical Summarization 3 Nov 2023 · 1 repository · arXiv:2311.02271Syntology ran 3 of 6 samples · 3 unverified
-
Abstractive Summarization of Large Document Collections Using GPT 9 Oct 2023 · 0 repositories · arXiv:2310.05690
-
Minimum-length chain embedding for the phase unwrapping problem on D-Wave's advantage architecture 19 Sep 2023 · 0 repositories · arXiv:2309.10296
-
Automatic Personalized Impression Generation for PET Reports Using Large Language Models 18 Sep 2023 · 2 repositories · arXiv:2309.10066
-
Multi-document Summarization: A Comparative Evaluation 10 Sep 2023 · 0 repositories · arXiv:2309.04951
-
"Beware of deception": Detecting Half-Truth and Debunking it through Controlled Claim Editing 15 Aug 2023 · 0 repositories · arXiv:2308.07973
-
Summarization from Leaderboards to Practice: Choosing A Representation Backbone and Ensuring Robustness 18 Jun 2023 · 0 repositories · arXiv:2306.10555
-
pyBibX -- A Python Library for Bibliometric and Scientometric Analysis Powered with Artificial Intelligence Tools 27 Apr 2023 · 1 repository · arXiv:2304.14516
-
Summaries as Captions: Generating Figure Captions for Scientific Documents with Automated Text Summarization 23 Feb 2023 · 0 repositories · arXiv:2302.12324
-
Unsupervised Summarization Re-ranking 19 Dec 2022 · 2 repositories · arXiv:2212.09593
-
Implementing Deep Learning-Based Approaches for Article Summarization in Indian Languages 12 Dec 2022 · 0 repositories · arXiv:2212.05702
-
Investigating Efficiently Extending Transformers for Long Input Summarization 8 Aug 2022 · 2 repositories · arXiv:2208.04347
-
Indian Legal Text Summarization: A Text Normalisation-based Approach 13 Jun 2022 · 0 repositories · arXiv:2206.06238
-
Medical Scientific Table-to-Text Generation with Human-in-the-Loop under the Data Sparsity Constraint 24 May 2022 · 0 repositories · arXiv:2205.12368
Tasks archive 2025-07-28
20 shown of 62 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections