{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/in-other-news-a-bi-style-text-to-speech-model","title":"In Other News: A Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data","arxiv_id":"1904.02790","date":"2019-04-04","proceeding":"NAACL 2019 6","authors":["Nishant Prateek","Mateusz Łajszczak","Roberto Barra-Chicote","Thomas Drugman","Jaime Lorenzo-Trueba","Thomas Merritt","Srikanth Ronanki","Trevor Wood"],"abstract":"Neural text-to-speech synthesis (NTTS) models have shown significant progress\nin generating high-quality speech, however they require a large quantity of\ntraining data. This makes creating models for multiple styles expensive and\ntime-consuming. In this paper different styles of speech are analysed based on\nprosodic variations, from this a model is proposed to synthesise speech in the\nstyle of a newscaster, with just a few hours of supplementary data. We pose the\nproblem of synthesising in a target style using limited data as that of\ncreating a bi-style model that can synthesise both neutral-style and\nnewscaster-style speech via a one-hot vector which factorises the two styles.\nWe also propose conditioning the model on contextual word embeddings, and\nextensively evaluate it against neutral NTTS, and neutral concatenative-based\nsynthesis. This model closes the gap in perceived style-appropriateness between\nnatural recordings for newscaster-style of speech, and neutral speech synthesis\nby approximately two-thirds.","url_abs":"http://arxiv.org/abs/1904.02790v1","url_pdf":"http://arxiv.org/pdf/1904.02790v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"in-other-news-a-bi-style-text-to-speech-model","repo_url":"https://github.com/inconnu11/Objective-evaluation_speech_synthesis","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"speech-synthesis","task_name":"Speech Synthesis"},{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"text-to-speech-synthesis","task_name":"Text-To-Speech Synthesis"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}