Papers › WaveNet: A Generative Model for Raw Audio
WaveNet: A Generative Model for Raw Audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, Koray Kavukcuoglu
This paper introduces WaveNet, a deep neural network for generating raw audio waveforms. The model is fully probabilistic and autoregressive, with the predictive distribution for each audio sample conditioned on all previous ones; nonetheless we show that it can be efficiently trained on data with tens of thousands of samples per second of audio. When applied to text-to-speech, it yields state-of-the-art performance, with human listeners rating it as significantly more natural sounding than the best parametric and concatenative systems for both English and Mandarin. A single WaveNet can capture the characteristics of many different speakers with equal fidelity, and can switch between them by conditioning on the speaker identity. When trained to model music, we find that it generates novel and often highly realistic musical fragments. We also show that it can be employed as a discriminative model, returning promising results for phoneme recognition.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="1609.03499")
Code
Syntology Ran 41 of 103 code samples harvested from 28 repositories linked to this paper; 62 have no recorded run. Of those that ran: 2 ran · honoured contract; 1 ran · violated contract; 7 ran · our draft was wrong; 5 ran · fixture could not drive it; 26 ran with no contract checked.
By repository: community (archive-listed): 103 samples from 28 repositories, 41 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.
62 repositories listed; official and paper-mentioned ones first.
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
103 samples harvested; 41 ran; 2 honoured the contract we drafted; 62 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.
Licence: 25 of the 103 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.
Harvested from 28 repositories linked to this paper, official or community; each sample names its own and says which. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.
Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.
Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.
9382ada555ac8131 · report
0d5b595bd813f5d2 · report
5170661e509fdac5 · report
79f4a42767095c68 · report
aeed3dc3c228cb85 · report
d48c7052e4342b63 · report
2a41b9a5b5561707 · report
3579d050a889047c · report
d792a0fba4201cf8 · report
e64b73bbdaad21c1 · report
d11e91181dcc9a07 · report
52e40ac0e567678c · report
177000aa16ed0914 · report
507c33200751b432 · report
ab75326a9401df22 · report
aa8c3981092c6563 · report
8816614f0f19222e · report
98dbe4ae7dc18a75 · report
a082de3ab11812a2 · report
4e167db5b9f6fc08 · report
09786ed59e7e9cf2 · report
65c6a6bb03ab096a · report
f89017533ce3657d · report
89203853d074962e · report
c45b9649408ac9fe · report
279056e98d5e0375 · report
6b4b9533fe39cda6 · report
72c1ea6847fe785b · report
3f851150ab8a5983 · report
4881eb48ec733168 · report
143bb37d58864fbb · report
98a4f58cb03b805d · report
c178be7285ad41cc · report
8a1ea8ba7390e065 · report
0007bb74d458f72c · report
204025577056e168 · report
4cc58ece006978a9 · report
5dede1b4064f5b66 · report
3ef824c52e18d496 · report
09179961802765aa · report
fa16967c039eb346 · report
12ebdce4997310a0 · report
f54d3bc20cbb0d72 · report
0a63ddb2755bae27 · report
97379a11b326f8df · report
763e43e907687af7 · report
2b41471e5bdd700e · report
481100beb31ff019 · report
f78382df6a2928e7 · report
7828c27785724f82 · report
7196b6c0ab860fa4 · report
1cded66daaa7d06f · report
9559930f1fe44156 · report
d7a9d958cef12e9b · report
95e029ea7951fceb · report
53a7436bad789a1b · report
827ad2a926b86c17 · report
cf9ba15ed1f4e6ce · report
b51cb3c367f37836 · report
7b74db674fb72fdf · report
12751d3cd265b98d · report
3a12c3886a1fc338 · report
a0b1b1b7e437c1aa · report
67b083d6b310c125 · report
10a70eeea6ab946f · report
64f12982e6a585b9 · report
9c3ef54612737736 · report
d128a002ad6e4cf7 · report
82626582e93e4865 · report
15a452673512283b · report
d8c9e892bf3ebab6 · report
40a217379db47f94 · report
24497247ed4e4626 · report
44d903b3f04aa30a · report
605a3067a4dbc087 · report
1b7121b9a228766d · report
931b3cb69e4b92b0 · report
2fae45e513f08b93 · report
3a87c4faeb7c34dc · report
d0e871c0709173b8 · report
6bd837a8733616b3 · report
5b9dfc24846c1bb1 · report
8ce9ce5e4bfcc589 · report
1b1b873049a95522 · report
9d6af79974e65de0 · report
9c5ee587758b40ea · report
757fe9a84dcc4315 · report
5d62167ff5d0c13e · report
01739949766dd095 · report
1214b92968fec5fb · report
7add735b85b4d35f · report
85002c58267da006 · report
693668a628d421fa · report
f86589e53de0342d · report
b004c6acf315cb1f · report
a08d0165dc5b2ab3 · report
36ca07d7d0f7af21 · report
e2e958ef5467a988 · report
3aad1f3fc6a3f8c4 · report
d082485663961b2a · report
5435df463cbd6b13 · report
915819bd9388e0b6 · report
00b8912699a3a152 · report
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Speech Synthesis | Mandarin Chinese | WaveNet (L+F) | Mean Opinion Score | 4.08 | #1 of 3 | Archive leaderboard | report |
| Speech Synthesis | Mandarin Chinese | LSTM-RNN parametric | Mean Opinion Score | 3.79 | #2 of 3 | Archive leaderboard | report |
| Speech Synthesis | Mandarin Chinese | HMM-driven concatenative | Mean Opinion Score | 3.47 | #3 of 3 | Archive leaderboard | report |
| Speech Synthesis | North American English | WaveNet (L+F) | Mean Opinion Score | 4.21 | #3 of 7 | Archive leaderboard | report |
| Speech Synthesis | North American English | HMM-driven concatenative | Mean Opinion Score | 3.86 | #5 of 7 | Archive leaderboard | report |
| Speech Synthesis | North American English | LSTM-RNN parametric | Mean Opinion Score | 3.67 | #6 of 7 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Introduced by this paper: Causal Convolution, Dilated Causal Convolution, WaveNet
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections