Browse State-of-the-Art › Text-To-Speech Synthesis › Papers, page 2
Text-To-Speech Synthesis
Papers archive 2025-07-28
archive papers tagged: 332 · with a code link: 104 · where Syntology ran a sample: 37 (35 with a run with no instrument failure, 2 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (37 of 332 tagged: 35 with a run with no instrument failure, 2 where every run was a failure of Syntology's instrument)
Page 2 of 4: papers 101 to 200 of 332, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
4 Apr 2019 1 repository listed
-
27 Mar 2019 1 repository listed
-
29 Oct 2018 1 repository listed
-
The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems25 Jun 2018 1 repository listed Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
S2ST-Omni: An Efficient and Scalable Multilingual Speech-to-Speech Translation Framework via Seamless Speech-Text Alignment and Streaming Speech Generation11 Jun 2025 0 repositories listed
-
A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions4 Jun 2025 0 repositories listed
-
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech3 Jun 2025 0 repositories listed
-
SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction2 Jun 2025 0 repositories listed
-
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems31 May 2025 0 repositories listed
-
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling26 May 2025 0 repositories listed
-
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis25 May 2025 0 repositories listed
-
FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation20 May 2025 0 repositories listed
-
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis18 May 2025 0 repositories listed
-
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications12 May 2025 0 repositories listed
-
A Multi-Agent Framework for Automated Qinqiang Opera Script Generation Using Large Language Models22 Apr 2025 0 repositories listed
-
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis14 Apr 2025 0 repositories listed
-
Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis14 Apr 2025 0 repositories listed
-
ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech13 Feb 2025 0 repositories listed
-
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron10 Jan 2025 0 repositories listed
-
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control10 Jan 2025 0 repositories listed
-
Probing Speaker-specific Features in Speaker Representations9 Jan 2025 0 repositories listed
-
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting28 Dec 2024 0 repositories listed
-
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis22 Dec 2024 0 repositories listed
-
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis16 Dec 2024 0 repositories listed
-
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens13 Dec 2024 0 repositories listed
-
Debatts: Zero-Shot Debating Text-to-Speech Synthesis10 Nov 2024 0 repositories listed
-
A Unified Framework for Collecting Text-to-Speech Synthesis Datasets for 22 Indian Languages18 Oct 2024 0 repositories listed
-
DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis17 Oct 2024 0 repositories listed
-
Bahasa Harmony: A Comprehensive Dataset for Bahasa Text-to-Speech Synthesis with Discrete Codec Modeling of EnGen-TTS9 Oct 2024 0 repositories listed
-
Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch9 Oct 2024 0 repositories listed
-
HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis6 Oct 2024 0 repositories listed
-
Generative Semantic Communication for Text-to-Speech Synthesis4 Oct 2024 0 repositories listed
-
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS30 Sep 2024 0 repositories listed
-
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis24 Sep 2024 0 repositories listed
-
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion16 Sep 2024 0 repositories listed
-
Text-To-Speech Synthesis In The Wild13 Sep 2024 0 repositories listed
-
Full-text Error Correction for Chinese Speech Recognition with Large Language Model12 Sep 2024 0 repositories listed
-
What happens to diffusion model likelihood when your model is conditional?10 Sep 2024 0 repositories listed
-
AS-Speech: Adaptive Style For Speech Synthesis9 Sep 2024 0 repositories listed
-
Speech Bandwidth Expansion Via High Fidelity Generative Adversarial Networks26 Jul 2024 0 repositories listed
-
Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models18 Jul 2024 0 repositories listed
-
Autoregressive Speech Synthesis without Vector Quantization11 Jul 2024 0 repositories listed
-
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis4 Jul 2024 0 repositories listed
-
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization2 Jul 2024 0 repositories listed
-
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis30 Jun 2024 0 repositories listed
-
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis16 Jun 2024 0 repositories listed
-
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment12 Jun 2024 0 repositories listed
-
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis8 Jun 2024 0 repositories listed
-
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers8 Jun 2024 0 repositories listed
-
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model6 Jun 2024 0 repositories listed
-
Style Mixture of Experts for Expressive Text-To-Speech Synthesis5 Jun 2024 0 repositories listed
-
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis4 Jun 2024 0 repositories listed
-
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback2 Jun 2024 0 repositories listed
-
DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models23 May 2024 0 repositories listed
-
Evaluating Text-to-Speech Synthesis from a Large Discrete Token-based Speech Language Model16 May 2024 0 repositories listed
-
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis4 Apr 2024 0 repositories listed
-
PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders3 Apr 2024 0 repositories listed
-
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters10 Jan 2024 0 repositories listed
-
Boosting Large Language Model for Speech Synthesis: An Empirical Study30 Dec 2023 0 repositories listed
-
Normalization of Lithuanian Text Using Regular Expressions29 Dec 2023 0 repositories listed
-
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis17 Dec 2023 0 repositories listed
-
An Experimental Study: Assessing the Combined Framework of WavLM and BEST-RQ for Text-to-Speech Synthesis8 Dec 2023 0 repositories listed
-
Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis6 Dec 2023 0 repositories listed
-
Code-Mixed Text to Speech Synthesis under Low-Resource Constraints2 Dec 2023 0 repositories listed
-
Guided Flows for Generative Modeling and Decision Making22 Nov 2023 0 repositories listed
-
Generative Pre-training for Speech with Flow Matching25 Oct 2023 0 repositories listed
-
8 Oct 2023 0 repositories listed
-
The VoiceMOS Challenge 2023: Zero-shot Subjective Speech Quality Prediction for Multiple Domains4 Oct 2023 0 repositories listed
-
DurIAN-E: Duration Informed Attention Network For Expressive Text-to-Speech Synthesis22 Sep 2023 0 repositories listed
-
The FruitShell French synthesis system at the Blizzard 2023 Challenge1 Sep 2023 0 repositories listed
-
Towards Spontaneous Style Modeling with Semi-supervised Pre-training for Conversational Text-to-Speech Synthesis31 Aug 2023 0 repositories listed
-
SALTTS: Leveraging Self-Supervised Speech Representations for improved Text-to-Speech Synthesis2 Aug 2023 0 repositories listed
-
Comparing normalizing flows and diffusion models for prosody and acoustic modelling in text-to-speech31 Jul 2023 0 repositories listed
-
SLMGAN: Exploiting Speech Language Model Representations for Unsupervised Zero-Shot Voice Conversion in GANs18 Jul 2023 0 repositories listed
-
High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units29 Jun 2023 0 repositories listed
-
ZET-Speech: Zero-shot adaptive Emotion-controllable Text-to-Speech Synthesis with Diffusion and Style-based Models23 May 2023 0 repositories listed
-
VAKTA-SETU: A Speech-to-Speech Machine Translation Service in Select Indic Languages21 May 2023 0 repositories listed
-
MParrotTTS: Multilingual Multi-speaker Text to Speech Synthesis in Low Resource Setting19 May 2023 0 repositories listed
-
A unified front-end framework for English text-to-speech synthesis18 May 2023 0 repositories listed
-
Accented Text-to-Speech Synthesis with Limited Data8 May 2023 0 repositories listed
-
M2-CTTS: End-to-End Multi-scale Multi-modal Conversational Text-to-Speech Synthesis3 May 2023 0 repositories listed
-
A Review of Deep Learning Techniques for Speech Processing30 Apr 2023 0 repositories listed
-
Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model24 Apr 2023 0 repositories listed
-
Text is All You Need: Personalizing ASR Models using Controllable Speech Synthesis27 Mar 2023 0 repositories listed
-
A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI23 Mar 2023 0 repositories listed
-
Controllable Prosody Generation With Partial Inputs14 Mar 2023 0 repositories listed
-
Do Prosody Transfer Models Transfer Prosody?7 Mar 2023 0 repositories listed
-
ParrotTTS: Text-to-Speech synthesis by exploiting self-supervised representations1 Mar 2023 0 repositories listed
-
UzbekTagger: The rule-based POS tagger for Uzbek language30 Jan 2023 0 repositories listed
-
Applying Automated Machine Translation to Educational Video Courses9 Jan 2023 0 repositories listed
-
ReVISE: Self-Supervised Speech Resynthesis With Visual Input for Universal and Generalized Speech Regeneration1 Jan 2023 0 repositories listed
-
21 Dec 2022 0 repositories listed
-
Investigation of Japanese PnG BERT language model in text-to-speech synthesis for pitch accent language16 Dec 2022 0 repositories listed
-
Text-to-speech synthesis based on latent variable conversion using diffusion probabilistic model and variational autoencoder16 Dec 2022 0 repositories listed
-
Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models17 Nov 2022 0 repositories listed
-
Technology Pipeline for Large Scale Cross-Lingual Dubbing of Lecture Videos into Multiple Indian Languages1 Nov 2022 0 repositories listed
-
Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-To-Speech27 Oct 2022 0 repositories listed
-
An Overview of Affective Speech Synthesis and Conversion in the Deep Learning Era6 Oct 2022 0 repositories listed
-
Controllable Accented Text-to-Speech Synthesis22 Sep 2022 0 repositories listed
-
EPIC TTS Models: Empirical Pruning Investigations Characterizing Text-To-Speech Models22 Sep 2022 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.