Papers › Transcending Scaling Laws with 0.1% Extra Compute

Transcending Scaling Laws with 0.1% Extra Compute

20 Oct 2022arXiv:2210.11399archive 2025-07-28

Yi Tay, Jason Wei, Hyung Won Chung, Vinh Q. Tran, David R. So, Siamak Shakeri, Xavier Garcia, Huaixiu Steven Zheng, Jinfeng Rao, Aakanksha Chowdhery, Denny Zhou, Donald Metzler, Slav Petrov, Neil Houlsby, Quoc V. Le, Mostafa Dehghani

Scaling language models improves performance but comes with significant computational costs. This paper proposes UL2R, a method that substantially improves existing language models and their scaling curves with a relatively tiny amount of extra compute. The key idea is to continue training a state-of-the-art large language model (e.g., PaLM) on a few more steps with UL2's mixture-of-denoiser objective. We show that, with almost negligible extra computational costs and no new sources of data, we are able to substantially improve the scaling properties of large language models on downstream metrics. In this paper, we continue training PaLM with UL2R, introducing a new set of models at 8B, 62B, and 540B scale which we call U-PaLM. Impressively, at 540B scale, we show an approximately 2x computational savings rate where U-PaLM achieves the same performance as the final PaLM 540B model at around half its computational budget (i.e., saving ∼4.4 million TPUv4 hours). We further show that this improved scaling curve leads to 'emergent abilities' on challenging BIG-Bench tasks -- for instance, U-PaLM does much better than PaLM on some tasks or demonstrates better quality at much smaller scale (62B as opposed to 540B). Overall, we show that U-PaLM outperforms PaLM on many few-shot setups, i.e., English NLP tasks (e.g., commonsense reasoning, question answering), reasoning tasks with chain-of-thought (e.g., GSM8K), multilingual tasks (MGSM, TydiQA), MMLU and challenging BIG-Bench tasks. Finally, we provide qualitative examples showing the new capabilities of U-PaLM for single and multi-span infilling.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Arithmetic ReasoningCross-Lingual Question AnsweringGSM8KLanguage ModellingLarge Language ModelMMLUMulti-task Language UnderstandingQuestion Answering

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Arithmetic Reasoning GSM8K U-PaLM Accuracy 58.5 #119 of 164 Archive leaderboard report
Arithmetic Reasoning GSM8K U-PaLM Parameters (Billion) 540 #119 of 164 Archive leaderboard report
Cross-Lingual Question Answering TyDiQA-GoldP U-PaLM 62B (fine-tuned) EM 78.4 #2 of 11 Archive leaderboard report
Cross-Lingual Question Answering TyDiQA-GoldP U-PaLM 62B (fine-tuned) F1 88.5 #2 of 11 Archive leaderboard report
Cross-Lingual Question Answering TyDiQA-GoldP U-PaLM-540B (CoT) EM 54.6 #6 of 11 Archive leaderboard report
Multi-task Language Understanding MGSM U-PaLM 540B (CoT) Average (%) 49.9 #7 of 12 Archive leaderboard report
Question Answering StrategyQA U-PaLM 540B Accuracy 76.6 #4 of 12 Archive leaderboard report
Question Answering StrategyQA PaLM 540B Accuracy 76.4 #5 of 12 Archive leaderboard report
Question Answering StrategyQA Minerva 540B Accuracy 61.9 #6 of 12 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

PaLM

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections