Papers › Accelerating Large Language Model Decoding with Speculative Sampling

Accelerating Large Language Model Decoding with Speculative Sampling

2 Feb 2023arXiv:2302.01318archive 2025-07-28

Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent SIfre, John Jumper

We present speculative sampling, an algorithm for accelerating transformer decoding by enabling the generation of multiple tokens from each transformer call. Our algorithm relies on the observation that the latency of parallel scoring of short continuations, generated by a faster but less powerful draft model, is comparable to that of sampling a single token from the larger target model. This is combined with a novel modified rejection sampling scheme which preserves the distribution of the target model within hardware numerics. We benchmark speculative sampling with Chinchilla, a 70 billion parameter language model, achieving a 2-2.5x decoding speedup in a distributed setup, without compromising the sample quality or making modifications to the model itself.

PaperPDFCode

In Syntology View this paper on Syntology: its repositories, every harvested function with whether it ran, its licence and the call to fetch it.

Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

apoorvumang/prompt-lookup-decoding mentioned on GitHubpytorch report
eth-sri/language-model-arithmetic mentioned on GitHubpytorchMIT report
feifeibear/llmspeculativesampling mentioned on GitHubpytorch report
nlpodyssey/rwkv.f90 mentioned on GitHubMIT report
wdrink/simplear mentioned on GitHubjaxMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Language ModelingLanguage ModellingLarge Language Modelmodel

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

Chinchilla

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections