{"url":"/method/pegasus","slug":"pegasus","name":"PEGASUS","full_name":"PEGASUS","full_name_withheld":false,"description_markdown":"**PEGASUS** proposes a transformer-based model for abstractive summarization. It uses a special self-supervised pre-training objective called gap-sentences generation (GSG) that's designed to perform well on summarization-related downstream tasks. As reported in the paper, \"both GSG and MLM are applied simultaneously to this example as pre-training objectives. Originally there are three sentences. One sentence is masked with [MASK1] and used as target generation text (GSG). The other two sentences remain in the input, but some tokens are randomly masked by [MASK2].\"","description_state":"present","introduced_year":null,"introduced_by":{"title":"PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization","paper":"/paper/pegasus-pre-training-with-extracted-gap","first_author":"Jingqing Zhang","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/pegasus-pre-training-with-extracted-gap"},"source":{"url":"https://arxiv.org/abs/1912.08777v2","title":"PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":53,"archive_num_papers":53,"papers_newest_first":[{"paper":null,"title":"Pegasus: A Universal Framework for Scalable Deep Learning Inference on the Dataplane","date":"2025-06-06","arxiv_id":"2506.05779","n_code_links":0,"syntology":null},{"paper":null,"title":"QUAD-LLM-MLTC: Large Language Models Ensemble Learning for Healthcare Text Multi-Label Classification","date":"2025-02-20","arxiv_id":"2502.14189","n_code_links":0,"syntology":null},{"paper":null,"title":"Implementing Large Quantum Boltzmann Machines as Generative AI Models for Dataset Balancing","date":"2025-02-05","arxiv_id":"2502.03086","n_code_links":0,"syntology":null},{"paper":null,"title":"Extract-and-Abstract: Unifying Extractive and Abstractive Summarization within Single Encoder-Decoder Framework","date":"2024-09-18","arxiv_id":"2409.11827","n_code_links":0,"syntology":null},{"paper":"/paper/glimmer-incorporating-graph-and-lexical","title":"GLIMMER: Incorporating Graph and Lexical Features in Unsupervised Multi-Document Summarization","date":"2024-08-19","arxiv_id":"2408.10115","n_code_links":1,"syntology":null},{"paper":"/paper/biolay-ak-ss-at-biolaysumm-domain-adaptation","title":"BioLay_AK_SS at BioLaySumm: Domain Adaptation by Two-Stage Fine-Tuning of Large Language Models used for Biomedical Lay Summary Generation","date":"2024-08-16","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Factual Dialogue Summarization via Learning from Large Language Models","date":"2024-06-20","arxiv_id":"2406.14709","n_code_links":0,"syntology":null},{"paper":null,"title":"Comparing Quantum Annealing and Spiking Neuromorphic Computing for Sampling Binary Sparse Coding QUBO Problems","date":"2024-05-30","arxiv_id":"2405.20525","n_code_links":0,"syntology":null},{"paper":null,"title":"Evaluating Text Summaries Generated by Large Language Models Using OpenAI's GPT","date":"2024-05-07","arxiv_id":"2405.04053","n_code_links":0,"syntology":null},{"paper":"/paper/medvoc-vocabulary-adaptation-for-fine-tuning","title":"MEDVOC: Vocabulary Adaptation for Fine-tuning Pre-trained Language Models on Medical Text Summarization","date":"2024-05-07","arxiv_id":"2405.04163","n_code_links":1,"syntology":{"ran":8,"of":9,"unverified":1,"pointer_only":9}},{"paper":null,"title":"Analysis of Multidomain Abstractive Summarization Using Salience Allocation","date":"2024-02-19","arxiv_id":"2402.11955","n_code_links":0,"syntology":null},{"paper":null,"title":"PEGASUS: Personalized Generative 3D Avatars with Composable Attributes","date":"2024-02-16","arxiv_id":"2402.10636","n_code_links":0,"syntology":null},{"paper":"/paper/source-identification-in-abstractive","title":"Source Identification in Abstractive Summarization","date":"2024-02-07","arxiv_id":"2402.04677","n_code_links":1,"syntology":null},{"paper":"/paper/pegasus-physically-enhanced-gaussian","title":"PEGASUS: Physically Enhanced Gaussian Splatting Simulation System for 6DoF Object Pose Dataset Generation","date":"2024-01-04","arxiv_id":"2401.02281","n_code_links":1,"syntology":null},{"paper":"/paper/revisiting-zero-shot-abstractive","title":"Revisiting Zero-Shot Abstractive Summarization in the Era of Large Language Models from the Perspective of Position Bias","date":"2024-01-03","arxiv_id":"2401.01989","n_code_links":1,"syntology":null},{"paper":"/paper/harnessing-the-power-of-prompt-based","title":"Harnessing the Power of Prompt-based Techniques for Generating School-Level Questions using Large Language Models","date":"2023-12-02","arxiv_id":"2312.01032","n_code_links":1,"syntology":null},{"paper":"/paper/famesumm-investigating-and-improving","title":"FaMeSumm: Investigating and Improving Faithfulness of Medical Summarization","date":"2023-11-03","arxiv_id":"2311.02271","n_code_links":1,"syntology":{"ran":3,"of":6,"unverified":3,"pointer_only":0}},{"paper":null,"title":"Abstractive Summarization of Large Document Collections Using GPT","date":"2023-10-09","arxiv_id":"2310.05690","n_code_links":0,"syntology":null},{"paper":null,"title":"Minimum-length chain embedding for the phase unwrapping problem on D-Wave's advantage architecture","date":"2023-09-19","arxiv_id":"2309.10296","n_code_links":0,"syntology":null},{"paper":"/paper/automatic-personalized-impression-generation","title":"Automatic Personalized Impression Generation for PET Reports Using Large Language Models","date":"2023-09-18","arxiv_id":"2309.10066","n_code_links":2,"syntology":null},{"paper":null,"title":"Multi-document Summarization: A Comparative Evaluation","date":"2023-09-10","arxiv_id":"2309.04951","n_code_links":0,"syntology":null},{"paper":null,"title":"\"Beware of deception\": Detecting Half-Truth and Debunking it through Controlled Claim Editing","date":"2023-08-15","arxiv_id":"2308.07973","n_code_links":0,"syntology":null},{"paper":null,"title":"Summarization from Leaderboards to Practice: Choosing A Representation Backbone and Ensuring Robustness","date":"2023-06-18","arxiv_id":"2306.10555","n_code_links":0,"syntology":null},{"paper":"/paper/pybibx-a-python-library-for-bibliometric-and","title":"pyBibX -- A Python Library for Bibliometric and Scientometric Analysis Powered with Artificial Intelligence Tools","date":"2023-04-27","arxiv_id":"2304.14516","n_code_links":1,"syntology":null},{"paper":null,"title":"Summaries as Captions: Generating Figure Captions for Scientific Documents with Automated Text Summarization","date":"2023-02-23","arxiv_id":"2302.12324","n_code_links":0,"syntology":null},{"paper":"/paper/unsupervised-summarization-re-ranking","title":"Unsupervised Summarization Re-ranking","date":"2022-12-19","arxiv_id":"2212.09593","n_code_links":2,"syntology":null},{"paper":null,"title":"Implementing Deep Learning-Based Approaches for Article Summarization in Indian Languages","date":"2022-12-12","arxiv_id":"2212.05702","n_code_links":0,"syntology":null},{"paper":"/paper/investigating-efficiently-extending","title":"Investigating Efficiently Extending Transformers for Long Input Summarization","date":"2022-08-08","arxiv_id":"2208.04347","n_code_links":2,"syntology":null},{"paper":null,"title":"Indian Legal Text Summarization: A Text Normalisation-based Approach","date":"2022-06-13","arxiv_id":"2206.06238","n_code_links":0,"syntology":null},{"paper":null,"title":"Medical Scientific Table-to-Text Generation with Human-in-the-Loop under the Data Sparsity Constraint","date":"2022-05-24","arxiv_id":"2205.12368","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/abstractive-text-summarization","name":"Abstractive Text Summarization","papers":22},{"task":"/task/text-summarization","name":"Text Summarization","papers":12},{"task":"/task/articles","name":"Articles","papers":6},{"task":"/task/document-summarization","name":"Document Summarization","papers":6},{"task":"/task/decoder","name":"Decoder","papers":5},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":4},{"task":"/task/multi-document-summarization","name":"Multi-Document Summarization","papers":4},{"task":"/task/sentence","name":"Sentence","papers":4},{"task":"/task/active-learning","name":"Active Learning","papers":3},{"task":"/task/domain-adaptation","name":"Domain Adaptation","papers":3},{"task":"/task/re-ranking","name":"Re-Ranking","papers":3},{"task":"/task/text-generation","name":"Text Generation","papers":3},{"task":"/task/bayesian-inference","name":"Bayesian Inference","papers":2},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":2},{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":2},{"task":"/task/language-modeling","name":"Language Modeling","papers":2},{"task":"/task/language-modelling","name":"Language Modelling","papers":2},{"task":"/task/mixture-of-experts","name":"Mixture-of-Experts","papers":2},{"task":"/task/question-answering","name":"Question Answering","papers":2},{"task":"/task/question-generation","name":"Question Generation","papers":2}],"tasks_shown":20,"n_tasks":62,"usage_by_year":[{"year":"2019","papers":1},{"year":"2020","papers":5},{"year":"2021","papers":16},{"year":"2022","papers":6},{"year":"2023","papers":10},{"year":"2024","papers":12},{"year":"2025","papers":3}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/pegasus"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}