{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/yourtts-towards-zero-shot-multi-speaker-tts","title":"YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone","arxiv_id":"2112.02418","date":"2021-12-04","proceeding":null,"authors":["Edresson Casanova","Julian Weber","Christopher Shulby","Arnaldo Candido Junior","Eren Gölge","Moacir Antonelli Ponti"],"abstract":"YourTTS brings the power of a multilingual approach to the task of zero-shot multi-speaker TTS. Our method builds upon the VITS model and adds several novel modifications for zero-shot multi-speaker and multilingual training. We achieved state-of-the-art (SOTA) results in zero-shot multi-speaker TTS and results comparable to SOTA in zero-shot voice conversion on the VCTK dataset. Additionally, our approach achieves promising results in a target language with a single-speaker dataset, opening possibilities for zero-shot multi-speaker TTS and zero-shot voice conversion systems in low-resource languages. Finally, it is possible to fine-tune the YourTTS model with less than 1 minute of speech and achieve state-of-the-art results in voice similarity and with reasonable quality. This is important to allow synthesis for speakers with a very different voice or recording characteristics from those seen during training.","url_abs":"https://arxiv.org/abs/2112.02418v4","url_pdf":"https://arxiv.org/pdf/2112.02418v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"yourtts-towards-zero-shot-multi-speaker-tts","repo_url":"https://github.com/coqui-ai/TTS","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MPL-2.0"}},{"paper_slug":"yourtts-towards-zero-shot-multi-speaker-tts","repo_url":"https://github.com/edresson/yourtts","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"yourtts-towards-zero-shot-multi-speaker-tts","repo_url":"https://github.com/daniilrobnikov/vits2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"speech-synthesis","task_name":"Speech Synthesis"},{"task_slug":"text-to-speech-synthesis","task_name":"Text-To-Speech Synthesis"},{"task_slug":"voice-conversion","task_name":"Voice Conversion"},{"task_slug":"voice-similarity","task_name":"Voice Similarity"},{"task_slug":"zero-shot-learning","task_name":"Zero-Shot Learning"},{"task_slug":"zero-shot-multi-speaker-tts","task_name":"Zero-Shot Multi-Speaker TTS"}],"methods":[{"method_slug":"hifi-gan","method_name":"HiFi-GAN"},{"method_slug":"normalizing-flows","method_name":"Normalizing Flows"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2112.02418","atlas_url":"https://app.syntology.ai/?focus=2112.02418","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}