{"url":"/method/melgan","slug":"melgan","name":"MelGAN","full_name":"MelGAN","full_name_withheld":false,"description_markdown":"**MelGAN** is a non-autoregressive feed-forward convolutional architecture to perform audio waveform generation in a [GAN](https://paperswithcode.com/method/gan) setup. The architecture is a fully convolutional feed-forward network with mel-spectrogram $s$ as input and raw waveform $x$ as output. Since the mel-spectrogram is at\r\na 256× lower temporal resolution, the authors use a stack of transposed convolutional layers to upsample the input sequence. Each transposed convolutional layer is followed by a stack of residual blocks with dilated convolutions. Unlike traditional GANs, the MelGAN generator does not use a global noise vector as input.\r\n\r\nTo deal with 'checkerboard artifacts' in audio, instead of using [PhaseShuffle](https://paperswithcode.com/method/phase-shuffle), MelGAN uses kernel-size as a multiple of stride.\r\n\r\n[Weight normalization](https://paperswithcode.com/method/weight-normalization) is used for normalization. A [window-based discriminator](https://paperswithcode.com/method/window-based-discriminator), similar to a [PatchGAN](https://paperswithcode.com/method/patchgan) is used for the discriminator.","description_state":"present","introduced_year":null,"introduced_by":{"title":"MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis","paper":"/paper/melgan-generative-adversarial-networks-for","first_author":"Kundan Kumar","n_authors":9,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/melgan-generative-adversarial-networks-for"},"source":{"url":"https://arxiv.org/abs/1910.06711v3","title":"MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Audio","area_id":"audio","collection":"Generative Audio Models","url":"/methods/category/generative-audio-models","pwc_aliases":[]}],"n_papers_tagged":14,"archive_num_papers":14,"papers_newest_first":[{"paper":null,"title":"WOLONet: Wave Outlooker for Efficient and High Fidelity Speech Synthesis","date":"2022-06-20","arxiv_id":"2206.09920","n_code_links":0,"syntology":null},{"paper":"/paper/cmelgan-an-efficient-conditional-generative","title":"cMelGAN: An Efficient Conditional Generative Model Based on Mel Spectrograms","date":"2022-05-15","arxiv_id":"2205.07319","n_code_links":1,"syntology":null},{"paper":"/paper/real-time-spectrogram-inversion-on-mobile","title":"Real time spectrogram inversion on mobile phone","date":"2022-03-01","arxiv_id":"2203.00756","n_code_links":1,"syntology":null},{"paper":null,"title":"Audio Deepfake Perceptions in College Going Populations","date":"2021-12-06","arxiv_id":"2112.03351","n_code_links":0,"syntology":null},{"paper":"/paper/vocbench-a-neural-vocoder-benchmark-for","title":"VocBench: A Neural Vocoder Benchmark for Speech Synthesis","date":"2021-12-06","arxiv_id":"2112.03099","n_code_links":1,"syntology":null},{"paper":null,"title":"Improve GAN-based Neural Vocoder using Pointwise Relativistic LeastSquare GAN","date":"2021-03-26","arxiv_id":"2103.14245","n_code_links":0,"syntology":null},{"paper":"/paper/tfgan-time-and-frequency-domain-based","title":"TFGAN: Time and Frequency Domain Based Generative Adversarial Network for High-fidelity Speech Synthesis","date":"2020-11-24","arxiv_id":"2011.12206","n_code_links":1,"syntology":null},{"paper":"/paper/universal-melgan-a-robust-neural-vocoder-for","title":"Universal MelGAN: A Robust Neural Vocoder for High-Fidelity Waveform Generation in Multiple Domains","date":"2020-11-19","arxiv_id":"2011.09631","n_code_links":2,"syntology":null},{"paper":"/paper/stylemelgan-an-efficient-high-fidelity","title":"StyleMelGAN: An Efficient High-Fidelity Adversarial Vocoder with Temporal Adaptive Normalization","date":"2020-11-03","arxiv_id":"2011.01557","n_code_links":2,"syntology":{"ran":0,"of":2,"unverified":2,"pointer_only":0}},{"paper":"/paper/speedyspeech-efficient-neural-speech","title":"SpeedySpeech: Efficient Neural Speech Synthesis","date":"2020-08-09","arxiv_id":"2008.03802","n_code_links":3,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}},{"paper":"/paper/vocgan-a-high-fidelity-real-time-vocoder-with","title":"VocGAN: A High-Fidelity Real-time Vocoder with a Hierarchically-nested Adversarial Network","date":"2020-07-30","arxiv_id":"2007.15256","n_code_links":2,"syntology":null},{"paper":"/paper/adversarial-representation-learning-for-2","title":"Adversarial representation learning for private speech generation","date":"2020-06-16","arxiv_id":"2006.09114","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":0}},{"paper":"/paper/se-melgan-speaker-agnostic-rapid-speech","title":"SE-MelGAN -- Speaker Agnostic Rapid Speech Enhancement","date":"2020-06-13","arxiv_id":"2006.07637","n_code_links":0,"syntology":null},{"paper":"/paper/melgan-generative-adversarial-networks-for","title":"MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis","date":"2019-10-08","arxiv_id":"1910.06711","n_code_links":21,"syntology":{"ran":5,"of":7,"unverified":2,"pointer_only":1}}],"papers_shown":14,"tasks":[{"task":"/task/speech-synthesis","name":"Speech Synthesis","papers":7},{"task":null,"name":"CPU","papers":6},{"task":null,"name":"GPU","papers":4},{"task":"/task/text-to-speech","name":"Text to Speech","papers":2},{"task":"/task/high","name":"Vocal Bursts Intensity Prediction","papers":2},{"task":"/task/text-to-speech-1","name":"text-to-speech","papers":2},{"task":"/task/audio-synthesis","name":"Audio Synthesis","papers":1},{"task":"/task/machine-learning","name":"BIG-bench Machine Learning","papers":1},{"task":"/task/face-swapping","name":"Face Swapping","papers":1},{"task":null,"name":"Generative Adversarial Network","papers":1},{"task":"/task/music-generation","name":"Music Generation","papers":1},{"task":"/task/privacy-preserving","name":"Privacy Preserving","papers":1},{"task":"/task/representation-learning","name":"Representation Learning","papers":1},{"task":"/task/spectral-reconstruction","name":"Spectral Reconstruction","papers":1},{"task":"/task/speech-enhancement","name":"Speech Enhancement","papers":1},{"task":"/task/translation","name":"Translation","papers":1}],"tasks_shown":16,"n_tasks":16,"usage_by_year":[{"year":"2019","papers":1},{"year":"2020","papers":7},{"year":"2021","papers":3},{"year":"2022","papers":3}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/melgan"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}