{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/voiceloop-voice-fitting-and-synthesis-via-a","title":"VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop","arxiv_id":"1707.06588","date":"2017-07-20","proceeding":"ICLR 2018 1","authors":["Yaniv Taigman","Lior Wolf","Adam Polyak","Eliya Nachmani"],"abstract":"We present a new neural text to speech (TTS) method that is able to transform\ntext to speech in voices that are sampled in the wild. Unlike other systems,\nour solution is able to deal with unconstrained voice samples and without\nrequiring aligned phonemes or linguistic features. The network architecture is\nsimpler than those in the existing literature and is based on a novel shifting\nbuffer working memory. The same buffer is used for estimating the attention,\ncomputing the output audio, and for updating the buffer itself. The input\nsentence is encoded using a context-free lookup table that contains one entry\nper character or phoneme. The speakers are similarly represented by a short\nvector that can also be fitted to new identities, even with only a few samples.\nVariability in the generated speech is achieved by priming the buffer prior to\ngenerating the audio. Experimental results on several datasets demonstrate\nconvincing capabilities, making TTS accessible to a wider range of\napplications. In order to promote reproducibility, we release our source code\nand models.","url_abs":"http://arxiv.org/abs/1707.06588v3","url_pdf":"http://arxiv.org/pdf/1707.06588v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"voiceloop-voice-fitting-and-synthesis-via-a","repo_url":"https://github.com/facebookresearch/loop","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"voiceloop-voice-fitting-and-synthesis-via-a","repo_url":"https://github.com/jasminsternkopf/mel_cepstral_distance","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1707.06588","atlas_url":"https://app.syntology.ai/?focus=1707.06588","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}