{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visualization-and-interpretation-of-latent","title":"Visualization and Interpretation of Latent Spaces for Controlling Expressive Speech Synthesis through Audio Analysis","arxiv_id":"1903.11570","date":"2019-03-27","proceeding":null,"authors":["Noé Tits","Fengna Wang","Kevin El Haddad","Vincent Pagel","Thierry Dutoit"],"abstract":"The field of Text-to-Speech has experienced huge improvements last years\nbenefiting from deep learning techniques. Producing realistic speech becomes\npossible now. As a consequence, the research on the control of the\nexpressiveness, allowing to generate speech in different styles or manners, has\nattracted increasing attention lately. Systems able to control style have been\ndeveloped and show impressive results. However the control parameters often\nconsist of latent variables and remain complex to interpret. In this paper, we\nanalyze and compare different latent spaces and obtain an interpretation of\ntheir influence on expressive speech. This will enable the possibility to build\ncontrollable speech synthesis systems with an understandable behaviour.","url_abs":"http://arxiv.org/abs/1903.11570v1","url_pdf":"http://arxiv.org/pdf/1903.11570v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visualization-and-interpretation-of-latent","repo_url":"https://github.com/noetits/ICE-Talk","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"emotional-speech-synthesis","task_name":"Emotional Speech Synthesis"},{"task_slug":"expressive-speech-synthesis","task_name":"Expressive Speech Synthesis"},{"task_slug":"learning-network-representations","task_name":"Learning Network Representations"},{"task_slug":"speech-emotion-recognition","task_name":"Speech Emotion Recognition"},{"task_slug":"speech-synthesis","task_name":"Speech Synthesis"},{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"text-to-speech-synthesis","task_name":"Text-To-Speech Synthesis"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.11570","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}