{"url":"/method/waveglow","slug":"waveglow","name":"WaveGlow","full_name":"WaveGlow","full_name_withheld":false,"description_markdown":"**WaveGlow** is a flow-based generative model that generates audio by sampling from a distribution. Specifically samples are taken from a zero mean spherical Gaussian with the same number of dimensions as our desired output, and those samples are put through a series of layers that transforms the simple distribution to one which has the desired distribution.","description_state":"present","introduced_year":null,"introduced_by":{"title":"WaveGlow: A Flow-based Generative Network for Speech Synthesis","paper":"/paper/waveglow-a-flow-based-generative-network-for","first_author":"Ryan Prenger","n_authors":3,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/waveglow-a-flow-based-generative-network-for"},"source":{"url":"http://arxiv.org/abs/1811.00002v1","title":"WaveGlow: A Flow-based Generative Network for Speech Synthesis","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Audio","area_id":"audio","collection":"Generative Audio Models","url":"/methods/category/generative-audio-models","pwc_aliases":[]}],"n_papers_tagged":21,"archive_num_papers":21,"papers_newest_first":[{"paper":null,"title":"Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach","date":"2024-09-10","arxiv_id":"2409.13734","n_code_links":0,"syntology":null},{"paper":null,"title":"Code-Mixed Text to Speech Synthesis under Low-Resource Constraints","date":"2023-12-02","arxiv_id":"2312.01103","n_code_links":0,"syntology":null},{"paper":null,"title":"Rapid Speaker Adaptation in Low Resource Text to Speech Systems using Synthetic Data and Transfer learning","date":"2023-12-02","arxiv_id":"2312.01107","n_code_links":0,"syntology":null},{"paper":null,"title":"Affective social anthropomorphic intelligent system","date":"2023-04-19","arxiv_id":"2304.11046","n_code_links":0,"syntology":null},{"paper":null,"title":"Adaptive re-calibration of channel-wise features for Adversarial Audio Classification","date":"2022-10-21","arxiv_id":"2210.11722","n_code_links":0,"syntology":null},{"paper":null,"title":"NatiQ: An End-to-end Text-to-Speech System for Arabic","date":"2022-06-15","arxiv_id":"2206.07373","n_code_links":0,"syntology":null},{"paper":null,"title":"FlowVocoder: A small Footprint Neural Vocoder based Normalizing flow for Speech Synthesis","date":"2021-09-27","arxiv_id":"2109.13675","n_code_links":0,"syntology":null},{"paper":"/paper/adaptation-of-tacotron2-based-text-to-speech","title":"Adaptation of Tacotron2-based Text-To-Speech for Articulatory-to-Acoustic Mapping using Ultrasound Tongue Imaging","date":"2021-07-26","arxiv_id":"2107.12051","n_code_links":1,"syntology":null},{"paper":null,"title":"A Flow-Based Neural Network for Time Domain Speech Enhancement","date":"2021-06-16","arxiv_id":"2106.09008","n_code_links":0,"syntology":null},{"paper":null,"title":"Low Bit-Rate Wideband Speech Coding: A Deep Generative Model based Approach","date":"2021-02-04","arxiv_id":"2102.02640","n_code_links":0,"syntology":null},{"paper":null,"title":"Text-to-speech for the hearing impaired","date":"2020-12-03","arxiv_id":"2012.02174","n_code_links":0,"syntology":null},{"paper":"/paper/melglow-efficient-waveform-generative-network","title":"MelGlow: Efficient Waveform Generative Network Based on Location-Variable Convolution","date":"2020-12-03","arxiv_id":"2012.01684","n_code_links":3,"syntology":null},{"paper":"/paper/stylemelgan-an-efficient-high-fidelity","title":"StyleMelGAN: An Efficient High-Fidelity Adversarial Vocoder with Temporal Adaptive Normalization","date":"2020-11-03","arxiv_id":"2011.01557","n_code_links":2,"syntology":{"ran":0,"of":2,"unverified":2,"pointer_only":0}},{"paper":null,"title":"A comparison of Vietnamese Statistical Parametric Speech Synthesis Systems","date":"2020-05-26","arxiv_id":"2005.12962","n_code_links":0,"syntology":null},{"paper":"/paper/probing-the-phonetic-and-phonological","title":"Probing the phonetic and phonological knowledge of tones in Mandarin TTS models","date":"2019-12-23","arxiv_id":"1912.10915","n_code_links":1,"syntology":null},{"paper":"/paper/waveflow-a-compact-flow-based-model-for-raw-1","title":"WaveFlow: A Compact Flow-based Model for Raw Audio","date":"2019-12-03","arxiv_id":"1912.01219","n_code_links":4,"syntology":{"ran":3,"of":9,"unverified":6,"pointer_only":0}},{"paper":null,"title":"Speaker independence of neural vocoders and their effect on parametric resynthesis speech enhancement","date":"2019-11-14","arxiv_id":"1911.06266","n_code_links":0,"syntology":null},{"paper":null,"title":"Transferring neural speech waveform synthesizers to musical instrument sounds generation","date":"2019-10-27","arxiv_id":"1910.12381","n_code_links":0,"syntology":null},{"paper":"/paper/parametric-resynthesis-with-neural-vocoders","title":"Parametric Resynthesis with neural vocoders","date":"2019-06-16","arxiv_id":"1906.06762","n_code_links":1,"syntology":null},{"paper":null,"title":"WaveCycleGAN2: Time-domain Neural Post-filter for Speech Waveform Generation","date":"2019-04-05","arxiv_id":"1904.02892","n_code_links":0,"syntology":null},{"paper":"/paper/waveglow-a-flow-based-generative-network-for","title":"WaveGlow: A Flow-based Generative Network for Speech Synthesis","date":"2018-10-31","arxiv_id":"1811.00002","n_code_links":2,"syntology":{"ran":2,"of":7,"unverified":5,"pointer_only":0}}],"papers_shown":21,"tasks":[{"task":"/task/speech-synthesis","name":"Speech Synthesis","papers":9},{"task":"/task/text-to-speech","name":"Text to Speech","papers":9},{"task":"/task/text-to-speech-1","name":"text-to-speech","papers":9},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":4},{"task":null,"name":"GPU","papers":3},{"task":"/task/audio-synthesis","name":"Audio Synthesis","papers":2},{"task":"/task/decoder","name":"Decoder","papers":2},{"task":"/task/density-estimation","name":"Density Estimation","papers":2},{"task":"/task/resynthesis","name":"Resynthesis","papers":2},{"task":"/task/speech-enhancement","name":"Speech Enhancement","papers":2},{"task":"/task/audio-classification","name":"Audio Classification","papers":1},{"task":"/task/audio-generation","name":"Audio Generation","papers":1},{"task":"/task/face-swapping","name":"Face Swapping","papers":1},{"task":"/task/quantization","name":"Quantization","papers":1},{"task":"/task/rhythm","name":"Rhythm","papers":1},{"task":"/task/spectral-reconstruction","name":"Spectral Reconstruction","papers":1},{"task":"/task/style-transfer","name":"Style Transfer","papers":1},{"task":"/task/synthetic-speech-detection","name":"Synthetic Speech Detection","papers":1},{"task":"/task/text-to-speech-synthesis","name":"Text-To-Speech Synthesis","papers":1},{"task":"/task/transliteration","name":"Transliteration","papers":1}],"tasks_shown":20,"n_tasks":23,"usage_by_year":[{"year":"2018","papers":1},{"year":"2019","papers":6},{"year":"2020","papers":4},{"year":"2021","papers":4},{"year":"2022","papers":2},{"year":"2023","papers":3},{"year":"2024","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/waveglow"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}