{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hierarchical-generative-modeling-for","title":"Hierarchical Generative Modeling for Controllable Speech Synthesis","arxiv_id":"1810.07217","date":"2018-10-16","proceeding":"ICLR 2019 5","authors":["Wei-Ning Hsu","Yu Zhang","Ron J. Weiss","Heiga Zen","Yonghui Wu","Yuxuan Wang","Yuan Cao","Ye Jia","Zhifeng Chen","Jonathan Shen","Patrick Nguyen","Ruoming Pang"],"abstract":"This paper proposes a neural sequence-to-sequence text-to-speech (TTS) model\nwhich can control latent attributes in the generated speech that are rarely\nannotated in the training data, such as speaking style, accent, background\nnoise, and recording conditions. The model is formulated as a conditional\ngenerative model based on the variational autoencoder (VAE) framework, with two\nlevels of hierarchical latent variables. The first level is a categorical\nvariable, which represents attribute groups (e.g. clean/noisy) and provides\ninterpretability. The second level, conditioned on the first, is a multivariate\nGaussian variable, which characterizes specific attribute configurations (e.g.\nnoise level, speaking rate) and enables disentangled fine-grained control over\nthese attributes. This amounts to using a Gaussian mixture model (GMM) for the\nlatent distribution. Extensive evaluation demonstrates its ability to control\nthe aforementioned attributes. In particular, we train a high-quality\ncontrollable TTS model on real found data, which is capable of inferring\nspeaker and style attributes from a noisy utterance and use it to synthesize\nclean speech with controllable speaking style.","url_abs":"http://arxiv.org/abs/1810.07217v2","url_pdf":"http://arxiv.org/pdf/1810.07217v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"hierarchical-generative-modeling-for","repo_url":"https://github.com/lturing/Tools","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"hierarchical-generative-modeling-for","repo_url":"https://github.com/rarefin/TTS_VAE","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"speech-synthesis","task_name":"Speech Synthesis"},{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1810.07217","atlas_url":"https://app.syntology.ai/?focus=1810.07217","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}