{"url":"/method/wavetts","slug":"wavetts","name":"WaveTTS","full_name":"WaveTTS","full_name_withheld":false,"description_markdown":"**WaveTTS** is a [Tacotron](https://paperswithcode.com/method/tacotron)-based text-to-speech architecture that has two loss functions: 1) time-domain loss, denoted as the waveform loss, that measures the distortion between the natural and generated waveform; and 2) frequency-domain loss, that measures the Mel-scale acoustic feature loss between the natural and generated acoustic features.\r\n\r\nThe motivation arises from [Tacotron 2](https://paperswithcode.com/method/tacotron-2). Here its feature prediction network is trained independently of the [WaveNet](https://paperswithcode.com/method/wavenet) vocoder. At run-time, the feature prediction network and WaveNet vocoder are artificially joined together. As a result, the framework suffers from the mismatch between frequency-domain acoustic features and time-domain waveform. To overcome such mismatch, WaveTTS uses a joint time-frequency domain loss for TTS that effectively improves the synthesized voice quality.","description_state":"present","introduced_year":null,"introduced_by":{"title":"WaveTTS: Tacotron-based TTS with Joint Time-Frequency Domain Loss","paper":"/paper/wavetts-tacotron-based-tts-with-joint-time","first_author":"Rui Liu","n_authors":5,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/wavetts-tacotron-based-tts-with-joint-time"},"source":{"url":"https://arxiv.org/abs/2002.00417v3","title":"WaveTTS: Tacotron-based TTS with Joint Time-Frequency Domain Loss","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Sequential","area_id":"sequential","collection":"Sequence To Sequence Models","url":"/methods/category/sequence-to-sequence-models","pwc_aliases":[]},{"area":"Audio","area_id":"audio","collection":"Text-to-Speech Models","url":"/methods/category/text-to-speech-models","pwc_aliases":[]}],"n_papers_tagged":1,"archive_num_papers":1,"papers_newest_first":[{"paper":"/paper/wavetts-tacotron-based-tts-with-joint-time","title":"WaveTTS: Tacotron-based TTS with Joint Time-Frequency Domain Loss","date":"2020-02-02","arxiv_id":"2002.00417","n_code_links":0,"syntology":null}],"papers_shown":1,"tasks":[{"task":"/task/text-to-speech","name":"Text to Speech","papers":1},{"task":"/task/text-to-speech-1","name":"text-to-speech","papers":1}],"tasks_shown":2,"n_tasks":2,"usage_by_year":[{"year":"2020","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/wavetts"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}