{"url":"/method/paranet","slug":"paranet","name":"ParaNet","full_name":"ParaNet","full_name_withheld":false,"description_markdown":"**ParaNet** is a non-autoregressive attention-based architecture for text-to-speech, which is fully convolutional and converts text to mel spectrogram. ParaNet distills the attention from the autoregressive text-to-spectrogram model, and iteratively refines the alignment between text and spectrogram in a layer-by-layer manner. The architecture is otherwise similar to [Deep Voice 3](https://paperswithcode.com/method/deep-voice-3) except these changes to the decoder; whereas the decoder of DV3 has multiple attention-based layers, where each layer consists of a\r\n[causal convolution](https://paperswithcode.com/method/causal-convolution) block followed by an attention block, ParaNet has a single attention block in the encoder.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/1905.08459v3","title":"Non-Autoregressive Neural Text-to-Speech","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Audio","area_id":"audio","collection":"Text-to-Speech Models","url":"/methods/category/text-to-speech-models","pwc_aliases":[]}],"n_papers_tagged":4,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"Estimating Parameters of the Tree Root in Heterogeneous Soil Environments via Mask-Guided Multi-Polarimetric Integration Neural Network","date":"2021-12-27","arxiv_id":"2112.13494","n_code_links":0,"syntology":null},{"paper":null,"title":"ParaNet: Deep Regular Representation for 3D Point Clouds","date":"2020-12-05","arxiv_id":"2012.03028","n_code_links":0,"syntology":null},{"paper":null,"title":"Parallel Neural Text-to-Speech","date":"2020-01-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/parallel-neural-text-to-speech","title":"Non-Autoregressive Neural Text-to-Speech","date":"2019-05-21","arxiv_id":"1905.08459","n_code_links":2,"syntology":{"ran":2,"of":7,"unverified":5,"pointer_only":0}}],"papers_shown":4,"tasks":[{"task":"/task/text-to-speech","name":"Text to Speech","papers":2},{"task":"/task/text-to-speech-1","name":"text-to-speech","papers":2},{"task":"/task/gpr","name":"GPR","papers":1},{"task":"/task/text-to-speech-synthesis","name":"Text-To-Speech Synthesis","papers":1},{"task":"/task/point-cloud-upsampling","name":"point cloud upsampling","papers":1}],"tasks_shown":5,"n_tasks":5,"usage_by_year":[{"year":"2019","papers":1},{"year":"2020","papers":2},{"year":"2021","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/paranet"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}