Methods › Sequential › Sequence To Sequence Models › Tacotron
Tacotron
Introduced by Yuxuan Wang et al. in Tacotron: Towards End-to-End Speech Synthesis
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Tacotron is an end-to-end generative text-to-speech model that takes a character sequence as input and outputs the corresponding spectrogram. The backbone of Tacotron is a seq2seq model with attention. The Figure depicts the model, which includes an encoder, an attention-based decoder, and a post-processing net. At a high-level, the model takes characters as input and produces spectrogram frames, which are then converted to waveforms.
Papers archive 2025-07-28
30 shown of 65, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech 29 Oct 2024 · 1 repository · arXiv:2410.22179
-
Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach 10 Sep 2024 · 0 repositories · arXiv:2409.13734
-
Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems 4 Sep 2024 · 0 repositories · arXiv:2409.02517
-
Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation 3 Apr 2024 · 0 repositories · arXiv:2404.02592
-
An overview of text-to-speech systems and media applications 22 Oct 2023 · 0 repositories · arXiv:2310.14301
-
Energy-Based Models For Speech Synthesis 19 Oct 2023 · 0 repositories · arXiv:2310.12765
-
The DeepZen Speech Synthesis System for Blizzard Challenge 2023 30 Aug 2023 · 0 repositories · arXiv:2308.15945
-
Multilingual Text-to-Speech Synthesis for Turkic Languages Using Transliteration 25 May 2023 · 1 repository · arXiv:2305.15749
-
A Virtual Simulation-Pilot Agent for Training of Air Traffic Controllers 16 Apr 2023 · 0 repositories · arXiv:2304.07842
-
ArmanTTS single-speaker Persian dataset 7 Apr 2023 · 0 repositories · arXiv:2304.03585
-
Investigation of Japanese PnG BERT language model in text-to-speech synthesis for pitch accent language 16 Dec 2022 · 0 repositories · arXiv:2212.08321
-
Investigating Content-Aware Neural Text-To-Speech MOS Prediction Using Prosodic and Linguistic Features 1 Nov 2022 · 0 repositories · arXiv:2211.00342
-
Cross-lingual Text-To-Speech with Flow-based Voice Conversion for Improved Pronunciation 31 Oct 2022 · 0 repositories · arXiv:2210.17264
-
Towards Developing State-of-the-Art TTS Synthesisers for 13 Indian Languages with Signal Processing aided Alignments 31 Oct 2022 · 0 repositories · arXiv:2210.17153
-
Efficiently Trained Low-Resource Mongolian Text-to-Speech System Based On FullConv-TTS 24 Oct 2022 · 0 repositories · arXiv:2211.01948
-
Facial Landmark Predictions with Applications to Metaverse 29 Sep 2022 · 1 repository · arXiv:2209.14698
-
Self-supervised learning for robust voice cloning 7 Apr 2022 · 0 repositories · arXiv:2204.03421
-
Singing-Tacotron: Global duration control attention and dynamic filter for End-to-end singing voice synthesis 16 Feb 2022 · 0 repositories · arXiv:2202.07907
-
Zero-Shot Long-Form Voice Cloning with Dynamic Convolution Attention 25 Jan 2022 · 0 repositories · arXiv:2201.10375
-
Word-Level Style Control for Expressive, Non-attentive Speech Synthesis 19 Nov 2021 · 0 repositories · arXiv:2111.10173
-
High Quality Streaming Speech Synthesis with Low, Sentence-Length-Independent Latency 17 Nov 2021 · 0 repositories · arXiv:2111.09052
-
On-device neural speech synthesis 17 Sep 2021 · 0 repositories · arXiv:2109.08710
-
Neural Sequence-to-Sequence Speech Synthesis Using a Hidden Semi-Markov Model Based Structured Attention Mechanism 31 Aug 2021 · 0 repositories · arXiv:2108.13985
-
Neural HMMs are all you need (for high-quality attention-free TTS) 30 Aug 2021 · 2 repositories · arXiv:2108.13320
-
One TTS Alignment To Rule Them All 23 Aug 2021 · 3 repositories · arXiv:2108.10447
-
Using Deep Learning Techniques and Inferential Speech Statistics for AI Synthesised Speech Recognition 23 Jul 2021 · 0 repositories · arXiv:2107.11412
-
AI based Presentation Creator With Customized Audio Content Delivery 27 Jun 2021 · 0 repositories · arXiv:2106.14213
-
Ctrl-P: Temporal Control of Prosodic Variation for Speech Synthesis 15 Jun 2021 · 0 repositories · arXiv:2106.08352
-
Exploring emotional prototypes in a high dimensional TTS latent space 5 May 2021 · 0 repositories · arXiv:2105.01891
-
VARA-TTS: Non-Autoregressive Text-to-Speech Synthesis based on Very Deep VAE with Residual Attention 12 Feb 2021 · 0 repositories · arXiv:2102.06431
Tasks archive 2025-07-28
20 shown of 49 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| Speech Synthesis | 43 |
| Text to Speech | 41 |
| text-to-speech | 41 |
| Text-To-Speech Synthesis | 15 |
| Decoder | 10 |
| Sentence | 6 |
| Transfer Learning | 5 |
| Voice Cloning | 5 |
| Speech Recognition | 4 |
| Voice Conversion | 4 |
| Audio Synthesis | 3 |
| Expressive Speech Synthesis | 3 |
| GPU | 3 |
| speech-recognition | 3 |
| All | 2 |
| CPU | 2 |
| Data Augmentation | 2 |
| Diversity | 2 |
| Generative Adversarial Network | 2 |
| Self-Supervised Learning | 2 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections