Browse State-of-the-Art › Text-To-Speech Synthesis
Text-To-Speech Synthesis
104 papers with code · 6 benchmarks · 18 datasets archive 2025-07-28
Text-To-Speech Synthesis is a machine learning task that involves converting written text into spoken words. The goal is to generate synthetic speech that sounds natural and resembles human speech as closely as possible.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| LJSpeech (16 rows) | NaturalSpeech | NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality | code | Syntology ran 0 of 3 samples · 3 unverified | Compare |
| 20000 utterances (1 row) | Mia | MIA-Prognosis: A Deep Learning Framework to Predict Therapy Response | code | — | Compare |
| CMUDict 0.7b (1 row) | Token-Level Ensemble Distillation | Token-Level Ensemble Distillation for Grapheme-to-Phoneme Conversion | — | — | Compare |
| HUI speech corpus (1 row) | Tacotron 2 | Neural Speech Synthesis in German | — | — | Compare |
| Thorsten voice 21.02 neutral (1 row) | Tacotron 2 | Neural Speech Synthesis in German | — | — | Compare |
| Trinity Speech-Gesture Dataset (1 row) | Match-TTSG | Unified speech and gesture synthesis using flow matching | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
18 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 104 papers with code (332 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
8 Jun 2020 37 repositories listed Syntology ran 73 of 119 samples · 46 unverified · 33 pointer-only (licence)In this paper, we propose FastSpeech 2, which addresses the issues in FastSpeech and better solves the one-to-many mapping problem in TTS by 1) directly training the model with ground-truth target instead of the…
-
29 Mar 2017 30 repositories listed Syntology ran 7 of 25 samples · 18 unverified · 6 pointer-only (licence)A text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module.
-
22 May 2019 22 repositories listed Syntology ran 3 of 11 samples · 8 unverified · 3 pointer-only (licence)In this work, we propose a novel feed-forward network based on Transformer to generate mel-spectrogram in parallel for TTS.
-
24 Oct 2017 22 repositories listed Syntology ran 1 of 28 samples · 27 unverified · 1 pointer-only (licence)This paper describes a novel text-to-speech (TTS) technique based on deep convolutional neural networks (CNN), without use of any recurrent units.
-
23 Feb 2018 16 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)The small number of weights in a Sparse WaveRNN makes it possible to sample high-fidelity audio on a mobile CPU in real time.
-
25 Oct 2019 12 repositories listed Syntology ran 0 of 20 samples · 20 unverified · 1 pointer-only (licence)We propose Parallel WaveGAN, a distillation-free, fast, and small-footprint waveform generation method using a generative adversarial network.
-
22 May 2019 11 repositories listedCompared with traditional concatenative and statistical parametric approaches, neural network based end-to-end models suffer from slow inference speed, and the synthesized speech is usually not robust (i.
-
12 Jun 2018 11 repositories listedClone a voice in 5 seconds to generate arbitrary speech in real-time
-
23 Mar 2018 11 repositories listed Syntology ran 6 of 21 samples · 15 unverified · 7 pointer-only (licence)In this work, we propose "global style tokens" (GSTs), a bank of embeddings that are jointly trained within Tacotron, a state-of-the-art end-to-end speech synthesis system.
-
6 May 2021 10 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 4 pointer-only (licence)Singing voice synthesis (SVS) systems are built to synthesize high-quality and expressive singing voice, in which the acoustic model generates the acoustic features (e.
-
5 Jan 2023 7 repositories listed Syntology ran 6 of 6 samples · 0 unverifiedIn addition, we find Vall-E could preserve the speaker's emotion and acoustic environment of the acoustic prompt in synthesis.
-
2 Sep 2020 7 repositories listed Syntology ran 1 of 3 samples · 2 unverified · 1 pointer-only (licence)This paper introduces WaveGrad, a conditional model for waveform generation which estimates gradients of the data density.
-
13 May 2021 6 repositories listed Syntology ran 9 of 18 samples · 9 unverified · 1 pointer-only (licence)Recently, denoising diffusion probabilistic models and generative score matching have shown high potential in modelling complex data distributions while stochastic calculus has provided a unified point of view on these…
-
22 May 2020 6 repositories listed Syntology ran 3 of 14 samples · 11 unverifiedBy leveraging the properties of flows, MAS searches for the most probable monotonic alignment between text and the latent representation of speech.
-
19 Sep 2018 6 repositories listedAlthough end-to-end neural text-to-speech (TTS) methods (such as Tacotron2) are proposed and achieve state-of-the-art performance, they still suffer from two problems: 1) low efficiency during training and inference; 2)…
-
4 Jun 2019 5 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)Capturing high-level structure in audio waveforms is challenging because a single second of audio spans tens of thousands of timesteps.
-
14 Jan 2019 5 repositories listedDuring the last few years, spoken language technologies have known a big improvement thanks to Deep Learning.
-
13 Jul 2022 4 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedThrough the preliminary study on diffusion model parameterization, we find that previous gradient-based TTS models require hundreds or thousands of iterations to guarantee high sample quality, which poses a challenge…
-
30 Sep 2021 4 repositories listed Syntology ran 6 of 12 samples · 6 unverified · 1 pointer-only (licence)Non-autoregressive text-to-speech (NAR-TTS) models such as FastSpeech 2 and Glow-TTS can synthesize high-quality speech from the given text in parallel.
-
15 Feb 2018 4 repositories listedIn this paper we introduce a set of resources and tools aimed at providing support for natural language processing, text-to-speech synthesis and speech recognition for Romanian.
-
18 Apr 2023 3 repositories listedKeywords: Bark, ai voice cloning, Suno, text-to-speech, artificial intelligence, audio generation, Meta's encodec, audio codebooks, semantic tokens, HuBert, transformer-based model, multilingual speech, wav2vec, linear…
-
31 May 2022 3 repositories listedWe develop machine translation and speech synthesis systems to complement the efforts of revitalizing Judeo-Spanish, the exiled language of Sephardic Jews, which survived for centuries, but now faces the threat of…
-
9 May 2022 3 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedIn this paper, we answer these questions by first defining the human-level quality based on the statistical significance of subjective measure and introducing appropriate guidelines to judge it, and then developing a…
-
4 Dec 2021 3 repositories listedYourTTS brings the power of a multilingual approach to the task of zero-shot multi-speaker TTS.
-
17 Jun 2021 3 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedThe model takes an input phoneme sequence, and through an iterative refinement process, generates an audio waveform.
-
15 Jun 2021 3 repositories listedIn order to meet the need for a high quality, publicly available male speech corpus within the field of speech recognition, we have designed and created RyanSpeech which contains textual materials from real-world…
-
12 May 2020 3 repositories listed Syntology ran 4 of 19 samples · 15 unverifiedIn this paper we propose Flowtron: an autoregressive flow-based generative network for text-to-speech synthesis with control over speech variation and style transfer.
-
25 Oct 2023 2 repositories listedThis paper proposes a method for investigating the impact of speech recognition errors on the performance of natural language understanding models.
-
7 Oct 2023 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedPrevious mainstream audio-and-text LLMs use discrete audio tokens to represent both input and output audio; however, they suffer from performance degradation on tasks such as automatic speech recognition, speech-to-text…
-
17 Nov 2022 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe open-source all models on the Bhashini platform.
Syntology lines on 20 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections