Methods › Audio › Generative Audio Models › WaveNet
WaveNet
Introduced by Aaron van den Oord et al. in WaveNet: A Generative Model for Raw Audio
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
WaveNet is an audio generative model based on the PixelCNN architecture. In order to deal with long-range temporal dependencies needed for raw audio generation, architectures are developed based on dilated causal convolutions, which exhibit very large receptive fields.
The joint probability of a waveform x⃗ = { x₁, …, x_T } is factorised as a product of conditional probabilities as follows:
p(x⃗) = ∏ₜ₌₁ᵀ p(xₜ |x₁, …,xₜ₋₁)
Each audio sample xₜ is therefore conditioned on the samples at all previous timesteps.
Papers archive 2025-07-28
30 shown of 171, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Aliasing Reduction in Neural Amp Modeling by Smoothing Activations 7 May 2025 · 0 repositories · arXiv:2505.04082
-
WaveNet-Volterra Neural Networks for Active Noise Control: A Fully Causal Approach 6 Apr 2025 · 1 repository · arXiv:2504.04450
-
An Ensemble Framework for Probabilistic Short-Term Load Forecasting Based on BiTCN and Deep Attention Networks 25 Feb 2025 · 1 repository
-
Explore the Use of Time Series Foundation Model for Car-Following Behavior Analysis 13 Jan 2025 · 0 repositories · arXiv:2501.07034
-
Autoregressive Speech Synthesis with Next-Distribution Prediction 22 Dec 2024 · 0 repositories · arXiv:2412.16846
-
SeagrassFinder: Deep Learning for Eelgrass Detection and Coverage Estimation in the Wild 20 Dec 2024 · 0 repositories · arXiv:2412.16147
-
Synthetic Time Series Data Generation for Healthcare Applications: A PCG Case Study 17 Dec 2024 · 0 repositories · arXiv:2412.16207
-
Deep Learning-Based Approach for Identification and Compensation of Nonlinear Distortions in Parametric Array Loudspeakers 2 Dec 2024 · 0 repositories · arXiv:2412.01092
-
Islanding Detection for Active Distribution Networks Using WaveNet+UNet Classifier 17 Oct 2024 · 0 repositories · arXiv:2410.13926
-
RF Challenge: The Data-Driven Radio Frequency Signal Separation Challenge 13 Sep 2024 · 1 repository · arXiv:2409.08839
-
InstructSing: High-Fidelity Singing Voice Generation via Instructing Yourself 10 Sep 2024 · 0 repositories · arXiv:2409.06330
-
Leveraging WaveNet for Dynamic Listening Head Modeling from Speech 8 Sep 2024 · 0 repositories · arXiv:2409.05089
-
Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems 4 Sep 2024 · 0 repositories · arXiv:2409.02517
-
Synthesizing Audio from Silent Video using Sequence to Sequence Modeling 25 Apr 2024 · 1 repository · arXiv:2404.17608
-
Foundational GPT Model for MEG 14 Apr 2024 · 1 repository · arXiv:2404.09256
-
A Novel Approach to WaveNet Architecture for RF Signal Separation with Learnable Dilation and Data Augmentation 8 Feb 2024 · 0 repositories · arXiv:2402.09461
-
Forecasting VIX using Bayesian Deep Learning 30 Jan 2024 · 0 repositories · arXiv:2401.17042
-
An overview of text-to-speech systems and media applications 22 Oct 2023 · 0 repositories · arXiv:2310.14301
-
Energy-Based Models For Speech Synthesis 19 Oct 2023 · 0 repositories · arXiv:2310.12765
-
WaveNet: Wave-Aware Image Enhancement 10 Oct 2023 · 1 repository
-
An Initial Exploration: Learning to Generate Realistic Audio for Silent Video 23 Aug 2023 · 1 repository · arXiv:2308.12408
-
Learning minimal representations of stochastic processes with variational autoencoders 21 Jul 2023 · 1 repository · arXiv:2307.11608
-
Speaker-independent neural formant synthesis 2 Jun 2023 · 0 repositories · arXiv:2306.01957
-
Multilingual Text-to-Speech Synthesis for Turkic Languages Using Transliteration 25 May 2023 · 1 repository · arXiv:2305.15749
-
Traffic Forecasting on New Roads Using Spatial Contrastive Pre-Training (SCPT) 9 May 2023 · 1 repository · arXiv:2305.05237
-
ArmanTTS single-speaker Persian dataset 7 Apr 2023 · 0 repositories · arXiv:2304.03585
-
Because Every Sensor Is Unique, so Is Every Pair: Handling Dynamicity in Traffic Forecasting 20 Feb 2023 · 1 repository · arXiv:2302.09956
-
Extreme Audio Time Stretching Using Neural Synthesis 30 Nov 2022 · 0 repositories · arXiv:2211.16992
-
WaveNets: Wavelet Channel Attention Networks 4 Nov 2022 · 1 repository · arXiv:2211.02695
-
Clarinet: A Music Retrieval System 23 Oct 2022 · 1 repository · arXiv:2210.12648
Tasks archive 2025-07-28
20 shown of 121 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| Speech Synthesis | 54 |
| Text to Speech | 51 |
| text-to-speech | 51 |
| Decoder | 18 |
| Text-To-Speech Synthesis | 16 |
| Voice Conversion | 12 |
| Audio Synthesis | 10 |
| Audio Generation | 8 |
| GPU | 8 |
| Time Series | 7 |
| Deep Learning | 6 |
| Speech Enhancement | 6 |
| Generative Adversarial Network | 5 |
| Speech Recognition | 5 |
| Time Series Analysis | 5 |
| Transfer Learning | 5 |
| Translation | 5 |
| CPU | 4 |
| Data Augmentation | 4 |
| Graph Neural Network | 4 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections