{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-variational-prosody-model-for-the","title":"A Variational Prosody Model for the decomposition and synthesis of speech prosody","arxiv_id":"1806.08685","date":"2018-06-22","proceeding":null,"authors":["Branislav Gerazov","Gérard Bailly","Omar Mohammed","Yi Xu","Philip N. Garner"],"abstract":"The quest for comprehensive generative models of intonation that link\nlinguistic and paralinguistic functions to prosodic forms has been a\nlongstanding challenge of speech communication research. More traditional\nintonation models have given way to the overwhelming performance of artificial\nintelligence (AI) techniques for training model-free, end-to-end mappings using\nmillions of tunable parameters. The shift towards machine learning models has\nnonetheless posed the reverse problem - a compelling need to discover\nknowledge, to explain, visualise and interpret. Our work bridges between a\ncomprehensive generative model of intonation and state-of-the-art AI\ntechniques. We build upon the modelling paradigm of the Superposition of\nFunctional Contours model and propose a Variational Prosody Model (VPM) that\nuses a network of deep variational contour generators to capture the\ncontext-sensitive variation of the constituent elementary prosodic cliches. We\nshow that the VPM can give insight into the intrinsic variability of these\nprosodic prototypes through learning a meaningful prosodic latent space\nrepresentation structure. We also show that the VPM brings improved modelling\nperformance especially when such variability is prominent. In a speech\nsynthesis scenario we believe the model can be used to generate a dynamic and\nnatural prosody contour largely devoid of averaging effects.","url_abs":"http://arxiv.org/abs/1806.08685v1","url_pdf":"http://arxiv.org/pdf/1806.08685v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-variational-prosody-model-for-the","repo_url":"https://github.com/gerazov/prosodeep","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"speech-synthesis","task_name":"Speech Synthesis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}