Papers › Bidirectional Variational Inference for Non-Autoregressive Text-to-Speech

Bidirectional Variational Inference for Non-Autoregressive Text-to-Speech

1 Jan 2021ICLR 2021 1archive 2025-07-28

Yoonhyung Lee, Joongbo Shin, Kyomin Jung

Although early text-to-speech (TTS) models such as Tacotron 2 have succeeded in generating human-like speech, their autoregressive (AR) architectures have a limitation that they require a lot of time to generate a mel-spectrogram consisting of hundreds of steps. In this paper, we propose a novel non-autoregressive TTS model called BVAE-TTS, which eliminates the architectural limitation and generates a mel-spectrogram in parallel. BVAE-TTS adopts bidirectional-inference variational autoencoder (BVAE) that learns hierarchical latent representations using both bottom-up and top-down paths to increase its expressiveness. To apply BVAE to TTS, we design our model to utilize text information via an attention mechanism. By using attention maps that BVAE-TTS generates, we train a duration predictor so that the model uses the predicted length of each phoneme at inference. In experiments conducted on LJSpeech dataset, we show that our model generates a mel-spectrogram 27 times faster than Tacotron 2 with similar speech quality. Furthermore, our BVAE-TTS outperforms Glow-TTS, which is one of the state-of-the-art non-autoregressive TTS models, in terms of both speech quality and inference speed while having 58% fewer parameters.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Text to SpeechVariational Inferencetext-to-speech

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

Activation NormalizationAffine CouplingBatch NormalizationBiGRUBiLSTMCBHGConvolutionDense ConnectionsDilated Causal ConvolutionDropoutGLOWGRUGlow-TTSGriffin-Lim AlgorithmHighway LayerHighway NetworkInvertible 1x1 ConvolutionLSTMLinear LayerLocation Sensitive AttentionMax PoolingMixture of Logistic DistributionsNormalizing FlowsReLUResidual ConnectionResidual GRUSigmoid ActivationTacotronTacotron 2Tanh ActivationWaveNetZoneout

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections