Papers › Muppet: Massive Multi-task Representations with Pre-Finetuning

Muppet: Massive Multi-task Representations with Pre-Finetuning

26 Jan 2021EMNLP 2021 11arXiv:2101.11038archive 2025-07-28

Armen Aghajanyan, Anchit Gupta, Akshat Shrivastava, Xilun Chen, Luke Zettlemoyer, Sonal Gupta

We propose pre-finetuning, an additional large-scale learning stage between language model pre-training and fine-tuning. Pre-finetuning is massively multi-task learning (around 50 datasets, over 4.8 million total labeled examples), and is designed to encourage learning of representations that generalize better to many different tasks. We show that pre-finetuning consistently improves performance for pretrained discriminators (e.g.~RoBERTa) and generation models (e.g.~BART) on a wide range of tasks (sentence prediction, commonsense reasoning, MRC, etc.), while also significantly improving sample efficiency during fine-tuning. We also show that large-scale multi-tasking is crucial; pre-finetuning can hurt performance when few tasks are used up until a critical point (usually above 15) after which performance improves linearly in the number of tasks.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Abstractive Text SummarizationCommon Sense ReasoningLanguage ModelingLanguage ModellingMulti-Task LearningNatural Language InferenceQuestion AnsweringSentenceSentence CompletionSentiment AnalysisText Summarization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Abstractive Text Summarization CNN / Daily Mail MUPPET BART Large ROUGE-1 44.45 #12 of 53 Archive leaderboard report
Abstractive Text Summarization CNN / Daily Mail MUPPET BART Large ROUGE-2 21.25 #12 of 53 Archive leaderboard report
Abstractive Text Summarization CNN / Daily Mail MUPPET BART Large ROUGE-L 41.4 #12 of 53 Archive leaderboard report
Common Sense Reasoning CommonsenseQA MUPPET Roberta Large Accuracy 79.2 #7 of 38 Archive leaderboard report
Natural Language Inference RTE MUPPET Roberta Large Accuracy 92.8% #6 of 90 Archive leaderboard report
Question Answering BoolQ MUPPET Roberta Large Accuracy 87.5 #13 of 65 Archive leaderboard report
Question Answering BoolQ MUPPET Roberta Base Accuracy 83.8 #20 of 65 Archive leaderboard report
Sentence Completion HellaSwag MUPPET Roberta Large Accuracy 86.4 #19 of 89 Archive leaderboard report
Sentiment Analysis SST-2 Binary classification MUPPET Roberta Large Accuracy 97.4 #4 of 87 Archive leaderboard report
Sentiment Analysis SST-2 Binary classification MUPPET Roberta base Accuracy 96.7 #12 of 87 Archive leaderboard report
Text Summarization GigaWord MUPPET BART Large ROUGE-1 40.4 #5 of 41 Archive leaderboard report
Text Summarization GigaWord MUPPET BART Large ROUGE-2 20.54 #5 of 41 Archive leaderboard report
Text Summarization GigaWord MUPPET BART Large ROUGE-L 36.21 #5 of 41 Archive leaderboard report
Text Summarization Reddit TIFU MUPPET BART Large ROUGE-1 30.3 #3 of 5 Archive leaderboard report
Text Summarization Reddit TIFU MUPPET BART Large ROUGE-2 11.25 #3 of 5 Archive leaderboard report
Text Summarization Reddit TIFU MUPPET BART Large ROUGE-L 24.92 #3 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections