Methods › Computer Vision › Generative Adversarial Networks › HiFi-GAN
HiFi-GAN
Introduced by Jungil Kong et al. in HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
HiFi-GAN is a generative adversarial network for speech synthesis. HiFi-GAN consists of one generator and two discriminators: multi-scale and multi-period discriminators. The generator and discriminators are trained adversarially, along with two additional losses for improving training stability and model performance.
The generator is a fully convolutional neural network. It uses a mel-spectrogram as input and upsamples it through transposed convolutions until the length of the output sequence matches the temporal resolution of raw waveforms. Every transposed convolution is followed by a multi-receptive field fusion (MRF) module.
For the discriminator, a multi-period discriminator (MPD) is used consisting of several sub-discriminators each handling a portion of periodic signals of input audio. Additionally, to capture consecutive patterns and long-term dependencies, the multi-scale discriminator (MSD) proposed in MelGAN is used, which consecutively evaluates audio samples at different levels.
Papers archive 2025-07-28
30 shown of 33, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
RingFormer: A Neural Vocoder with Ring Attention and Convolution-Augmented Transformer 2 Jan 2025 · 1 repository · arXiv:2501.01182
-
A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction 11 Dec 2024 · 0 repositories · arXiv:2412.08312
-
TSELM: Target Speaker Extraction using Discrete Tokens and Language Models 12 Sep 2024 · 1 repository · arXiv:2409.07841
-
DSP-informed bandwidth extension using locally-conditioned excitation and linear time-varying filter subnetworks 22 Jul 2024 · 0 repositories · arXiv:2407.15624
-
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning 5 Jun 2024 · 1 repository · arXiv:2406.03049Syntology ran 4 of 4 samples · 0 unverified
-
CoVoMix: Advancing Zero-Shot Speech Generation for Human-like Multi-talker Conversations 10 Apr 2024 · 1 repository · arXiv:2404.06690
-
SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis 30 Jan 2024 · 0 repositories · arXiv:2402.01753
-
UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization 26 Jan 2024 · 0 repositories · arXiv:2401.14664
-
Scaling NVIDIA's Multi-speaker Multi-lingual TTS Systems with Zero-Shot TTS to Indic Languages 24 Jan 2024 · 0 repositories · arXiv:2401.13851
-
SELM: Speech Enhancement Using Discrete Tokens and Language Models 15 Dec 2023 · 0 repositories · arXiv:2312.09747
-
APNet2: High-quality and High-efficiency Neural Vocoder with Direct Prediction of Amplitude and Phase Spectra 20 Nov 2023 · 1 repository · arXiv:2311.11545
-
Collaborative Watermarking for Adversarial Speech Synthesis 26 Sep 2023 · 0 repositories · arXiv:2309.15224
-
HiFTNet: A Fast High-Quality Neural Vocoder with Harmonic-plus-Noise Filter and Inverse Short Time Fourier Transform 18 Sep 2023 · 1 repository · arXiv:2309.09493
-
Rep2wav: Noise Robust text-to-speech Using self-supervised representations 28 Aug 2023 · 0 repositories · arXiv:2308.14553
-
MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies 3 Aug 2023 · 1 repository · arXiv:2308.01546Syntology ran 9 of 14 samples · 5 unverified · 14 pointer-only (licence)
-
Speaker-independent neural formant synthesis 2 Jun 2023 · 0 repositories · arXiv:2306.01957
-
Source-Filter-Based Generative Adversarial Neural Vocoder for High Fidelity Speech Synthesis 26 Apr 2023 · 1 repository · arXiv:2304.13270
-
Wave-U-Net Discriminator: Fast and Lightweight Discriminator for Generative Adversarial Network-Based Speech Synthesis 24 Mar 2023 · 0 repositories · arXiv:2303.13909
-
Self-Supervised Representations for Singing Voice Conversion 21 Mar 2023 · 0 repositories · arXiv:2303.12197
-
Fast and small footprint Hybrid HMM-HiFiGAN based system for speech synthesis in Indian languages 13 Feb 2023 · 0 repositories · arXiv:2302.06227
-
MnTTS2: An Open-Source Multi-Speaker Mongolian Text-to-Speech Synthesis Dataset 11 Dec 2022 · 1 repository · arXiv:2301.00657
-
Hiding speaker's sex in speech using zero-evidence speaker representation in an analysis/synthesis pipeline 29 Nov 2022 · 1 repository · arXiv:2211.16065
-
Towards Building Text-To-Speech Systems for the Next Billion Users 17 Nov 2022 · 2 repositories · arXiv:2211.09536Syntology ran 0 of 1 samples · 1 unverified
-
Source-Filter HiFi-GAN: Fast and Pitch Controllable High-Fidelity Neural Vocoder 27 Oct 2022 · 0 repositories · arXiv:2210.15533
-
MnTTS: An Open-Source Mongolian Text-to-Speech Synthesis Dataset and Accompanied Baseline 22 Sep 2022 · 1 repository · arXiv:2209.10848
-
Mandarin Singing Voice Synthesis with Denoising Diffusion Probabilistic Wasserstein GAN 21 Sep 2022 · 0 repositories · arXiv:2209.10446
-
JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech 31 Mar 2022 · 2 repositories · arXiv:2203.16852Syntology ran 0 of 2 samples · 2 unverified
-
Disentangleing Content and Fine-grained Prosody Information via Hybrid ASR Bottleneck Features for Voice Conversion 24 Mar 2022 · 0 repositories · arXiv:2203.12813
-
ECAPA-TDNN for Multi-speaker Text-to-speech Synthesis 20 Mar 2022 · 1 repository · arXiv:2203.10473
-
iSTFTNet: Fast and Lightweight Mel-Spectrogram Vocoder Incorporating Inverse Short-Time Fourier Transform 4 Mar 2022 · 2 repositories · arXiv:2203.02395Syntology ran 8 of 11 samples · 3 unverified
Tasks archive 2025-07-28
20 shown of 61 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| Speech Synthesis | 16 |
| Text to Speech | 13 |
| text-to-speech | 13 |
| Text-To-Speech Synthesis | 7 |
| Voice Conversion | 6 |
| Generative Adversarial Network | 5 |
| Audio Generation | 3 |
| Decoder | 3 |
| Automatic Speech Recognition (ASR) | 2 |
| CPU | 2 |
| Self-Supervised Learning | 2 |
| Speaker Verification | 2 |
| Speech Enhancement | 2 |
| Speech Recognition | 2 |
| Voice Cloning | 2 |
| speech-recognition | 2 |
| Automatic Speech Recognition | 1 |
| Bandwidth Extension | 1 |
| Beat Tracking | 1 |
| Data Augmentation | 1 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections