{"url":"/method/deep-voice-3","slug":"deep-voice-3","name":"Deep Voice 3","full_name":"Deep Voice 3","full_name_withheld":false,"description_markdown":"**Deep Voice 3 (DV3)** is a fully-convolutional attention-based neural text-to-speech system. The Deep Voice 3 architecture consists of three components:\r\n\r\n- Encoder: A fully-convolutional encoder, which converts textual features to an internal\r\nlearned representation.\r\n\r\n- Decoder: A fully-convolutional causal decoder, which decodes the learned representation\r\nwith a multi-hop convolutional attention mechanism into a low-dimensional audio representation (mel-scale spectrograms) in an autoregressive manner.\r\n\r\n- Converter: A fully-convolutional post-processing network, which predicts final vocoder\r\nparameters (depending on the vocoder choice) from the decoder hidden states. Unlike the\r\ndecoder, the converter is non-causal and can thus depend on future context information.\r\n\r\nThe overall objective function to be optimized is a linear combination of the losses from the decoder and the converter. The authors separate decoder and converter and apply multi-task training, because it makes attention learning easier in practice. To be specific, the loss for mel-spectrogram prediction guides training of the attention mechanism, because the attention is trained with the gradients from mel-spectrogram prediction besides vocoder parameter prediction.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"http://arxiv.org/abs/1710.07654v3","title":"Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Audio","area_id":"audio","collection":"Text-to-Speech Models","url":"/methods/category/text-to-speech-models","pwc_aliases":[]}],"n_papers_tagged":3,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"Parallel Neural Text-to-Speech","date":"2020-01-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/parallel-neural-text-to-speech","title":"Non-Autoregressive Neural Text-to-Speech","date":"2019-05-21","arxiv_id":"1905.08459","n_code_links":2,"syntology":{"ran":6,"of":7,"unverified":1,"pointer_only":0}},{"paper":"/paper/deep-voice-3-scaling-text-to-speech-with","title":"Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning","date":"2017-10-20","arxiv_id":"1710.07654","n_code_links":7,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":1}}],"papers_shown":3,"tasks":[{"task":"/task/text-to-speech","name":"Text to Speech","papers":3},{"task":"/task/text-to-speech-1","name":"text-to-speech","papers":3},{"task":null,"name":"GPU","papers":1},{"task":"/task/speech-synthesis","name":"Speech Synthesis","papers":1},{"task":"/task/text-to-speech-synthesis","name":"Text-To-Speech Synthesis","papers":1}],"tasks_shown":5,"n_tasks":5,"usage_by_year":[{"year":"2017","papers":1},{"year":"2019","papers":1},{"year":"2020","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/deep-voice-3"},"syntology_read_at":"2026-09-25T09:33:49+00:00"}