{"url":"/method/fastspeech-2","slug":"fastspeech-2","name":"FastSpeech 2","full_name":"FastSpeech 2","full_name_withheld":false,"description_markdown":"**FastSpeech2** is a text-to-speech model that aims to improve upon FastSpeech by better solving the one-to-many mapping problem in TTS, i.e., multiple speech variations corresponding to the same text. It attempts to solve this problem by 1) directly training the model with ground-truth target instead of the simplified output from teacher, and 2) introducing more variation information of speech (e.g., pitch, energy and more accurate duration) as conditional inputs. Specifically, in FastSpeech 2, we extract duration, pitch and energy from speech waveform and directly take them as conditional inputs in training and use predicted values in inference.\r\n\r\nThe encoder converts the phoneme embedding sequence into the phoneme hidden sequence, and then the variance adaptor adds different variance information such as duration, pitch and energy into the hidden sequence, finally the mel-spectrogram decoder converts the adapted hidden sequence into mel-spectrogram sequence in parallel. FastSpeech 2 uses a feed-forward [Transformer](https://paperswithcode.com/method/transformer) block, which is a stack of [self-attention](https://paperswithcode.com/method/multi-head-attention) and 1D-[convolution](https://paperswithcode.com/method/convolution) as in FastSpeech, as the basic structure for the encoder and mel-spectrogram decoder.","description_state":"present","introduced_year":null,"introduced_by":{"title":"FastSpeech 2: Fast and High-Quality End-to-End Text to Speech","paper":"/paper/fastspeech-2-fast-and-high-quality-end-to-end","first_author":"Yi Ren","n_authors":7,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/fastspeech-2-fast-and-high-quality-end-to-end"},"source":{"url":"https://arxiv.org/abs/2006.04558v8","title":"FastSpeech 2: Fast and High-Quality End-to-End Text to Speech","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Audio","area_id":"audio","collection":"Text-to-Speech Models","url":"/methods/category/text-to-speech-models","pwc_aliases":[]}],"n_papers_tagged":20,"archive_num_papers":20,"papers_newest_first":[{"paper":null,"title":"AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis","date":"2025-04-12","arxiv_id":"2504.09225","n_code_links":0,"syntology":null},{"paper":null,"title":"AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation","date":"2024-12-13","arxiv_id":"2412.10103","n_code_links":0,"syntology":null},{"paper":null,"title":"Training Universal Vocoders with Feature Smoothing-Based Augmentation Methods for High-Quality TTS Systems","date":"2024-09-04","arxiv_id":"2409.02517","n_code_links":0,"syntology":null},{"paper":null,"title":"AttentionStitch: How Attention Solves the Speech Editing Problem","date":"2024-03-05","arxiv_id":"2403.04804","n_code_links":0,"syntology":null},{"paper":"/paper/back-transcription-as-a-method-for-evaluating","title":"Back Transcription as a Method for Evaluating Robustness of Natural Language Understanding Models to Speech Recognition Errors","date":"2023-10-25","arxiv_id":"2310.16609","n_code_links":2,"syntology":null},{"paper":null,"title":"Energy-Based Models For Speech Synthesis","date":"2023-10-19","arxiv_id":"2310.12765","n_code_links":0,"syntology":null},{"paper":"/paper/daspeech-directed-acyclic-transformer-for-1","title":"DASpeech: Directed Acyclic Transformer for Fast and High-quality Speech-to-Speech Translation","date":"2023-10-11","arxiv_id":"2310.07403","n_code_links":1,"syntology":{"ran":5,"of":7,"unverified":2,"pointer_only":7}},{"paper":"/paper/towards-robust-fastspeech-2-by-modelling","title":"Towards Robust FastSpeech 2 by Modelling Residual Multimodality","date":"2023-06-02","arxiv_id":"2306.01442","n_code_links":1,"syntology":null},{"paper":null,"title":"The Effects of Input Type and Pronunciation Dictionary Usage in Transfer Learning for Low-Resource Text-to-Speech","date":"2023-06-01","arxiv_id":"2306.00535","n_code_links":0,"syntology":null},{"paper":"/paper/libris2s-a-german-english-speech-to-speech","title":"LibriS2S: A German-English Speech-to-Speech Translation Corpus","date":"2022-04-22","arxiv_id":"2204.10593","n_code_links":1,"syntology":null},{"paper":null,"title":"Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech","date":"2022-03-31","arxiv_id":"2203.17190","n_code_links":0,"syntology":null},{"paper":"/paper/ecapa-tdnn-for-multi-speaker-text-to-speech","title":"ECAPA-TDNN for Multi-speaker Text-to-speech Synthesis","date":"2022-03-20","arxiv_id":"2203.10473","n_code_links":1,"syntology":null},{"paper":"/paper/multi-singer-fast-multi-singer-singing-voice-1","title":"Multi-Singer: Fast Multi-Singer Singing Voice Vocoder With A Large-Scale Corpus","date":"2021-12-20","arxiv_id":"2112.10358","n_code_links":2,"syntology":{"ran":0,"of":2,"unverified":2,"pointer_only":0}},{"paper":null,"title":"Improving Prosody for Unseen Texts in Speech Synthesis by Utilizing Linguistic Information and Noisy Data","date":"2021-11-15","arxiv_id":"2111.07549","n_code_links":0,"syntology":null},{"paper":"/paper/portaspeech-portable-and-high-quality","title":"PortaSpeech: Portable and High-Quality Generative Text-to-Speech","date":"2021-09-30","arxiv_id":"2109.15166","n_code_links":4,"syntology":{"ran":6,"of":12,"unverified":6,"pointer_only":1}},{"paper":"/paper/one-tts-alignment-to-rule-them-all","title":"One TTS Alignment To Rule Them All","date":"2021-08-23","arxiv_id":"2108.10447","n_code_links":3,"syntology":null},{"paper":null,"title":"Digital Einstein Experience: Fast Text-to-Speech for Conversational AI","date":"2021-07-21","arxiv_id":"2107.10658","n_code_links":0,"syntology":null},{"paper":"/paper/investigating-on-incorporating-pretrained-and","title":"Investigating on Incorporating Pretrained and Learnable Speaker Representations for Multi-Speaker Multi-Style Text-to-Speech","date":"2021-03-06","arxiv_id":"2103.04088","n_code_links":1,"syntology":null},{"paper":null,"title":"Parallel waveform synthesis based on generative adversarial networks with voicing-aware conditional discriminators","date":"2020-10-27","arxiv_id":"2010.14151","n_code_links":0,"syntology":null},{"paper":"/paper/fastspeech-2-fast-and-high-quality-end-to-end","title":"FastSpeech 2: Fast and High-Quality End-to-End Text to Speech","date":"2020-06-08","arxiv_id":"2006.04558","n_code_links":37,"syntology":{"ran":73,"of":119,"unverified":46,"pointer_only":33}}],"papers_shown":20,"tasks":[{"task":"/task/text-to-speech","name":"Text to Speech","papers":13},{"task":"/task/text-to-speech-1","name":"text-to-speech","papers":13},{"task":"/task/speech-synthesis","name":"Speech Synthesis","papers":7},{"task":"/task/text-to-speech-synthesis","name":"Text-To-Speech Synthesis","papers":5},{"task":"/task/decoder","name":"Decoder","papers":2},{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":2},{"task":"/task/speech-to-speech-translation","name":"Speech-to-Speech Translation","papers":2},{"task":"/task/speech-to-text","name":"Speech-to-Text","papers":2},{"task":"/task/translation","name":"Translation","papers":2},{"task":"/task/all","name":"All","papers":1},{"task":"/task/audio-generation","name":"Audio Generation","papers":1},{"task":"/task/chinese-word-segmentation","name":"Chinese Word Segmentation","papers":1},{"task":"/task/cross-lingual-transfer","name":"Cross-Lingual Transfer","papers":1},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":1},{"task":"/task/multi-task-learning","name":"Multi-Task Learning","papers":1},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":1},{"task":"/task/pos","name":"POS","papers":1},{"task":"/task/pos-tagging","name":"POS Tagging","papers":1},{"task":"/task/part-of-speech-tagging","name":"Part-Of-Speech Tagging","papers":1},{"task":"/task/polyphone-disambiguation","name":"Polyphone disambiguation","papers":1}],"tasks_shown":20,"n_tasks":36,"usage_by_year":[{"year":"2020","papers":2},{"year":"2021","papers":6},{"year":"2022","papers":3},{"year":"2023","papers":5},{"year":"2024","papers":3},{"year":"2025","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/fastspeech-2"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}