{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-fully-time-domain-neural-model-for-subband","title":"A Fully Time-domain Neural Model for Subband-based Speech Synthesizer","arxiv_id":"1810.05319","date":"2018-10-12","proceeding":null,"authors":["Azam Rabiee","Soo-Young Lee"],"abstract":"This paper introduces a deep neural network model for subband-based speech\nsynthesizer. The model benefits from the short bandwidth of the subband signals\nto reduce the complexity of the time-domain speech generator. We employed the\nmulti-level wavelet analysis/synthesis to decompose/reconstruct the signal to\nsubbands in time domain. Inspired from the WaveNet, a convolutional neural\nnetwork (CNN) model predicts subband speech signals fully in time domain. Due\nto the short bandwidth of the subbands, a simple network architecture is enough\nto train the simple patterns of the subbands accurately. In the ground truth\nexperiments with teacher forcing, the subband synthesizer outperforms the\nfullband model significantly. In addition, by conditioning the model on the\nphoneme sequence using a pronunciation dictionary, we have achieved the first\nfully time-domain neural text-to-speech (TTS) system. The generated speech of\nthe subband TTS shows comparable quality as the fullband one with a slighter\nnetwork architecture for each subband.","url_abs":"http://arxiv.org/abs/1810.05319v1","url_pdf":"http://arxiv.org/pdf/1810.05319v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-fully-time-domain-neural-model-for-subband","repo_url":"https://github.com/AzamRabiee/subband-TTS","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}