Browse State-of-the-Art › text-to-speech
text-to-speech
395 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 395 papers with code (1,413 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
12 Sep 2016 62 repositories listed Syntology ran 41 of 103 samples · 62 unverified · 25 pointer-only (licence)This paper introduces WaveNet, a deep neural network for generating raw audio waveforms.
-
8 Jun 2020 37 repositories listed Syntology ran 73 of 119 samples · 46 unverified · 33 pointer-only (licence)In this paper, we propose FastSpeech 2, which addresses the issues in FastSpeech and better solves the one-to-many mapping problem in TTS by 1) directly training the model with ground-truth target instead of the…
-
29 Mar 2017 30 repositories listed Syntology ran 7 of 25 samples · 18 unverified · 6 pointer-only (licence)A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module.
-
22 May 2019 22 repositories listed Syntology ran 3 of 11 samples · 8 unverified · 3 pointer-only (licence)In this work, we propose a novel feed-forward network based on Transformer to generate mel-spectrogram in parallel for TTS.
-
24 Oct 2017 22 repositories listed Syntology ran 1 of 28 samples · 27 unverified · 1 pointer-only (licence)This paper describes a novel text-to-speech (TTS) technique based on deep convolutional neural networks (CNN), without use of any recurrent units.
-
23 Feb 2018 16 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)The small number of weights in a Sparse WaveRNN makes it possible to sample high-fidelity audio on a mobile CPU in real time.
-
25 Oct 2019 12 repositories listed Syntology ran 0 of 20 samples · 20 unverified · 1 pointer-only (licence)We propose Parallel WaveGAN, a distillation-free, fast, and small-footprint waveform generation method using a generative adversarial network.
-
22 May 2019 11 repositories listedCompared with traditional concatenative and statistical parametric approaches, neural network based end-to-end models suffer from slow inference speed, and the synthesized speech is usually not robust (i.
-
12 Jun 2018 11 repositories listedClone a voice in 5 seconds to generate arbitrary speech in real-time
-
6 May 2021 10 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 4 pointer-only (licence)Singing voice synthesis (SVS) systems are built to synthesize high-quality and expressive singing voice, in which the acoustic model generates the acoustic features (e.
-
15 Jun 2021 9 repositories listed Syntology ran 8 of 12 samples · 4 unverified · 1 pointer-only (licence)Using full-band mel-spectrograms as input, we expect to generate high-resolution signals by adding a discriminator that employs spectrograms of multiple resolutions as the input.
-
15 Nov 2018 8 repositories listedThis paper introduces a robust universal neural vocoder trained with 74 speakers (comprised of both genders) coming from 17 languages.
-
5 Jan 2023 7 repositories listed Syntology ran 6 of 6 samples · 0 unverifiedIn addition, we find Vall-E could preserve the speaker's emotion and acoustic environment of the acoustic prompt in synthesis.
-
20 Oct 2017 7 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We present Deep Voice 3, a fully-convolutional attention-based neural text-to-speech (TTS) system.
-
14 Nov 2012 7 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedOne of the key challenges in sequence transduction is learning to represent both the input and output sequences in a way that is invariant to sequential distortions such as shrinking, stretching and translating.
-
7 Jul 2021 6 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 3 pointer-only (licence)We present SoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs.
-
13 May 2021 6 repositories listed Syntology ran 9 of 18 samples · 9 unverified · 1 pointer-only (licence)Recently, denoising diffusion probabilistic models and generative score matching have shown high potential in modelling complex data distributions while stochastic calculus has provided a unified point of view on these…
-
8 Oct 2020 6 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedThis paper presents Non-Attentive Tacotron based on the Tacotron 2 text-to-speech model, replacing the attention mechanism with an explicit duration predictor.
-
11 Jun 2020 6 repositories listed Syntology ran 4 of 9 samples · 5 unverifiedWe present FastPitch, a fully-parallel text-to-speech model based on FastSpeech, conditioned on fundamental frequency contours.
-
5 Jun 2020 6 repositories listedTo explore this issue, we proposed to employ Mockingjay, a self-supervised learning based model, to protect anti-spoofing models against adversarial attacks in the black-box scenario.
-
22 May 2020 6 repositories listed Syntology ran 3 of 14 samples · 11 unverifiedBy leveraging the properties of flows, MAS searches for the most probable monotonic alignment between text and the latent representation of speech.
-
19 Sep 2018 6 repositories listedAlthough end-to-end neural text-to-speech (TTS) methods (such as Tacotron2) are proposed and achieve state-of-the-art performance, they still suffer from two problems: 1) low efficiency during training and inference; 2)…
-
4 Jun 2019 5 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)Capturing high-level structure in audio waveforms is challenging because a single second of audio spans tens of thousands of timesteps.
-
19 Jul 2018 5 repositories listed Syntology ran 1 of 13 samples · 12 unverifiedIn this work, we propose a new solution for parallel wave generation by WaveNet.
-
23 Sep 2017 5 repositories listedIn the proposed framework incorporating the GANs, the discriminator is trained to distinguish natural and generated speech parameters, while the acoustic models are trained to minimize the weighted sum of the…
-
22 Aug 2023 4 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages?
-
13 Jul 2022 4 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedThrough the preliminary study on diffusion model parameterization, we find that previous gradient-based TTS models require hundreds or thousands of iterations to guarantee high sample quality, which poses a challenge…
-
30 Sep 2021 4 repositories listed Syntology ran 6 of 12 samples · 6 unverified · 1 pointer-only (licence)Non-autoregressive text-to-speech (NAR-TTS) models such as FastSpeech 2 and Glow-TTS can synthesize high-quality speech from the given text in parallel.
-
14 Sep 2021 4 repositories listedThis paper presents fairseq S^2, a fairseq extension for speech synthesis.
-
8 Feb 2021 4 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedText to speech (TTS) has been broadly used to synthesize natural and intelligible speech in different scenarios.
Syntology lines on 23 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections