{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/stft-spectral-loss-for-training-a-neural","title":"STFT spectral loss for training a neural speech waveform model","arxiv_id":"1810.11945","date":"2018-10-29","proceeding":null,"authors":["Shinji Takaki","Toru Nakashika","Xin Wang","Junichi Yamagishi"],"abstract":"This paper proposes a new loss using short-time Fourier transform (STFT)\nspectra for the aim of training a high-performance neural speech waveform model\nthat predicts raw continuous speech waveform samples directly. Not only\namplitude spectra but also phase spectra obtained from generated speech\nwaveforms are used to calculate the proposed loss. We also mathematically show\nthat training of the waveform model on the basis of the proposed loss can be\ninterpreted as maximum likelihood training that assumes the amplitude and phase\nspectra of generated speech waveforms following Gaussian and von Mises\ndistributions, respectively. Furthermore, this paper presents a simple network\narchitecture as the speech waveform model, which is composed of uni-directional\nlong short-term memories (LSTMs) and an auto-regressive structure. Experimental\nresults showed that the proposed neural model synthesized high-quality speech\nwaveforms.","url_abs":"http://arxiv.org/abs/1810.11945v2","url_pdf":"http://arxiv.org/pdf/1810.11945v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"stft-spectral-loss-for-training-a-neural","repo_url":"https://github.com/nii-yamagishilab/TSNetVocoder","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}