Browse State-of-the-Art › text-to-speech › Papers, page 9
text-to-speech
Papers archive 2025-07-28
archive papers tagged: 1,413 · with a code link: 395 · where Syntology ran a sample: 106 (95 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (106 of 1,413 tagged: 95 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument)
Page 9 of 15: papers 801 to 900 of 1,413, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
MuLanTTS: The Microsoft Speech Synthesis System for Blizzard Challenge 20236 Sep 2023 0 repositories listed
-
PromptTTS 2: Describing and Generating Voices with Text Prompt5 Sep 2023 0 repositories listed
-
A Comparative Analysis of Pretrained Language Models for Text-to-Speech4 Sep 2023 0 repositories listed
-
Learning Speech Representation From Contrastive Token-Acoustic Pretraining1 Sep 2023 0 repositories listed
-
The FruitShell French synthesis system at the Blizzard 2023 Challenge1 Sep 2023 0 repositories listed
-
Improving Mandarin Prosodic Structure Prediction with Multi-level Contextual Information31 Aug 2023 0 repositories listed
-
Towards Spontaneous Style Modeling with Semi-supervised Pre-training for Conversational Text-to-Speech Synthesis31 Aug 2023 0 repositories listed
-
The DeepZen Speech Synthesis System for Blizzard Challenge 202330 Aug 2023 0 repositories listed
-
Pruning Self-Attention for Zero-Shot Multi-Speaker Text-to-Speech28 Aug 2023 0 repositories listed
-
Rep2wav: Noise Robust text-to-speech Using self-supervised representations28 Aug 2023 0 repositories listed
-
Generalizable Zero-Shot Speaker Adaptive Speech Synthesis with Disentangled Representations24 Aug 2023 0 repositories listed
-
Multi-GradSpeech: Towards Diffusion-based Multi-Speaker Text-to-speech Using Consistent Diffusion Models21 Aug 2023 0 repositories listed
-
AffectEcho: Speaker Independent and Language-Agnostic Emotion and Affect Transfer for Speech Synthesis16 Aug 2023 0 repositories listed
-
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer14 Aug 2023 0 repositories listed
-
SALTTS: Leveraging Self-Supervised Speech Representations for improved Text-to-Speech Synthesis2 Aug 2023 0 repositories listed
-
Comparing normalizing flows and diffusion models for prosody and acoustic modelling in text-to-speech31 Jul 2023 0 repositories listed
-
Improving grapheme-to-phoneme conversion by learning pronunciations from speech recordings31 Jul 2023 0 repositories listed
-
Multilingual context-based pronunciation learning for Text-to-Speech31 Jul 2023 0 repositories listed
-
METTS: Multilingual Emotional Text-to-Speech by Cross-speaker and Cross-lingual Emotion Transfer29 Jul 2023 0 repositories listed
-
Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic Coding28 Jul 2023 0 repositories listed
-
SLMGAN: Exploiting Speech Language Model Representations for Unsupervised Zero-Shot Voice Conversion in GANs18 Jul 2023 0 repositories listed
-
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis14 Jul 2023 0 repositories listed
-
Controllable Emphasis with zero data for text-to-speech13 Jul 2023 0 repositories listed
-
On the Use of Self-Supervised Speech Representations in Spontaneous Speech Synthesis11 Jul 2023 0 repositories listed
-
Artificial Eye for the Blind7 Jul 2023 0 repositories listed
-
ContextSpeech: Expressive and Efficient Text-to-Speech for Paragraph Reading3 Jul 2023 0 repositories listed
-
High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units29 Jun 2023 0 repositories listed
-
GenerTTS: Pronunciation Disentanglement for Timbre and Style Generalization in Cross-Lingual Text-to-Speech27 Jun 2023 0 repositories listed
-
DSE-TTS: Dual Speaker Embedding for Cross-Lingual Text-to-Speech25 Jun 2023 0 repositories listed
-
Visual-Aware Text-to-Speech21 Jun 2023 0 repositories listed
-
Expressive Machine Dubbing Through Phrase-level Cross-lingual Prosody Transfer20 Jun 2023 0 repositories listed
-
CML-TTS A Multilingual Dataset for Speech Synthesis in Low-Resource Languages16 Jun 2023 0 repositories listed
-
Low-Resource Text-to-Speech Using Specific Data and Noise Augmentation16 Jun 2023 0 repositories listed
-
Improving Code-Switching and Named Entity Recognition in ASR with Speech Editing based Data Augmentation14 Jun 2023 0 repositories listed
-
PauseSpeech: Natural Speech Synthesis via Pre-trained Language Model and Pause-based Prosody Modeling13 Jun 2023 0 repositories listed
-
Learning Emotional Representations from Imbalanced Speech Data for Speech Emotion Recognition and Emotional Text-to-Speech9 Jun 2023 0 repositories listed
-
Ada-TTA: Towards Adaptive High-Quality Text-to-Talking Avatar Synthesis6 Jun 2023 0 repositories listed
-
Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias6 Jun 2023 0 repositories listed
-
Cross-Lingual Transfer Learning for Phrase Break Prediction with Multilingual Language Model5 Jun 2023 0 repositories listed
-
Rhythm-controllable Attention with High Robustness for Long Sentence Speech Synthesis5 Jun 2023 0 repositories listed
-
The Effects of Input Type and Pronunciation Dictionary Usage in Transfer Learning for Low-Resource Text-to-Speech1 Jun 2023 0 repositories listed
-
Text-to-Speech Pipeline for Swiss German -- A comparison31 May 2023 0 repositories listed
-
LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus30 May 2023 0 repositories listed
-
Make-A-Voice: Unified Voice Synthesis With Discrete Representation30 May 2023 0 repositories listed
-
Resource-Efficient Fine-Tuning Strategies for Automatic MOS Prediction in Text-to-Speech for Low-Resource Languages30 May 2023 0 repositories listed
-
STT4SG-350: A Speech Corpus for All Swiss German Dialect Regions30 May 2023 0 repositories listed
-
Towards Selection of Text-to-speech Data to Augment ASR Training30 May 2023 0 repositories listed
-
Automatic Evaluation of Turn-taking Cues in Conversational Speech Synthesis29 May 2023 0 repositories listed
-
DisfluencyFixer: A tool to enhance Language Learning through Speech To Speech Disfluency Correction26 May 2023 0 repositories listed
-
VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation25 May 2023 0 repositories listed
-
LAraBench: Benchmarking Arabic AI with Large Language Models24 May 2023 0 repositories listed
-
ZET-Speech: Zero-shot adaptive Emotion-controllable Text-to-Speech Synthesis with Diffusion and Style-based Models23 May 2023 0 repositories listed
-
Text Generation with Speech Synthesis for ASR Data Augmentation22 May 2023 0 repositories listed
-
ViT-TTS: Visual Text-to-Speech with Scalable Diffusion Transformer22 May 2023 0 repositories listed
-
VAKTA-SETU: A Speech-to-Speech Machine Translation Service in Select Indic Languages21 May 2023 0 repositories listed
-
ComedicSpeech: Text To Speech For Stand-up Comedies in Low-Resource Scenarios20 May 2023 0 repositories listed
-
MParrotTTS: Multilingual Multi-speaker Text to Speech Synthesis in Low Resource Setting19 May 2023 0 repositories listed
-
A unified front-end framework for English text-to-speech synthesis18 May 2023 0 repositories listed
-
Data Redaction from Conditional Generative Models18 May 2023 0 repositories listed
-
FastFit: Towards Real-Time Iterative Neural Vocoder by Replacing U-Net Encoder With Multiple STFTs18 May 2023 0 repositories listed
-
Controllable Speaking Styles Using a Large Language Model17 May 2023 0 repositories listed
-
Accented Text-to-Speech Synthesis with Limited Data8 May 2023 0 repositories listed
-
M2-CTTS: End-to-End Multi-scale Multi-modal Conversational Text-to-Speech Synthesis3 May 2023 0 repositories listed
-
A Review of Deep Learning Techniques for Speech Processing30 Apr 2023 0 repositories listed
-
Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model24 Apr 2023 0 repositories listed
-
DiffVoice: Text-to-Speech with Latent Diffusion23 Apr 2023 0 repositories listed
-
A Virtual Simulation-Pilot Agent for Training of Air Traffic Controllers16 Apr 2023 0 repositories listed
-
Enhancing Speech-to-Speech Translation with Multiple TTS Targets10 Apr 2023 0 repositories listed
-
ArmanTTS single-speaker Persian dataset7 Apr 2023 0 repositories listed
-
Ensemble prosody prediction for expressive speech synthesis3 Apr 2023 0 repositories listed
-
Text is All You Need: Personalizing ASR Models using Controllable Speech Synthesis27 Mar 2023 0 repositories listed
-
Wave-U-Net Discriminator: Fast and Lightweight Discriminator for Generative Adversarial Network-Based Speech Synthesis24 Mar 2023 0 repositories listed
-
A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI23 Mar 2023 0 repositories listed
-
Code-Switching Text Generation and Injection in Mandarin-English ASR20 Mar 2023 0 repositories listed
-
Cross-speaker Emotion Transfer by Manipulating Speech Style Latents15 Mar 2023 0 repositories listed
-
Controllable Prosody Generation With Partial Inputs14 Mar 2023 0 repositories listed
-
QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis14 Mar 2023 0 repositories listed
-
An End-to-End Neural Network for Image-to-Audio Transformation10 Mar 2023 0 repositories listed
-
Text-to-ECG: 12-Lead Electrocardiogram Synthesis conditioned on Clinical Text Reports9 Mar 2023 0 repositories listed
-
Do Prosody Transfer Models Transfer Prosody?7 Mar 2023 0 repositories listed
-
FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model6 Mar 2023 0 repositories listed
-
Fine-grained Emotional Control of Text-To-Speech: Learning To Rank Inter- And Intra-Class Emotion Intensities2 Mar 2023 0 repositories listed
-
Leveraging Large Text Corpora for End-to-End Speech Summarization2 Mar 2023 0 repositories listed
-
LiteG2P: A fast, light and high accuracy model for grapheme-to-phoneme conversion2 Mar 2023 0 repositories listed
-
DTW-SiameseNet: Dynamic Time Warped Siamese Network for Mispronunciation Detection and Correction1 Mar 2023 0 repositories listed
-
ParrotTTS: Text-to-Speech synthesis by exploiting self-supervised representations1 Mar 2023 0 repositories listed
-
Automatic Heteronym Resolution Pipeline Using RAD-TTS Aligners28 Feb 2023 0 repositories listed
-
ClArTTS: An Open-Source Classical Arabic Text-to-Speech Corpus28 Feb 2023 0 repositories listed
-
CrossSpeech: Speaker-independent Acoustic Representation for Cross-lingual Speech Synthesis28 Feb 2023 0 repositories listed
-
UniFLG: Unified Facial Landmark Generator from Text or Speech28 Feb 2023 0 repositories listed
-
Duration-aware pause insertion using pre-trained language model for multi-speaker text-to-speech27 Feb 2023 0 repositories listed
-
Varianceflow: High-Quality and Controllable Text-to-Speech using Variance Information via Normalizing Flow27 Feb 2023 0 repositories listed
-
Emphasizing Unseen Words: New Vocabulary Acquisition for End-to-End Speech Recognition20 Feb 2023 0 repositories listed
-
Fast and small footprint Hybrid HMM-HiFiGAN based system for speech synthesis in Indian languages13 Feb 2023 0 repositories listed
-
MAC: A unified framework boosting low resource automatic speech recognition5 Feb 2023 0 repositories listed
-
UzbekTagger: The rule-based POS tagger for Uzbek language30 Jan 2023 0 repositories listed
-
On granularity of prosodic representations in expressive text-to-speech26 Jan 2023 0 repositories listed
-
Modelling low-resource accents without accent-specific TTS frontend11 Jan 2023 0 repositories listed
-
UnifySpeech: A Unified Framework for Zero-shot Text-to-Speech and Voice Conversion10 Jan 2023 0 repositories listed
-
Applying Automated Machine Translation to Educational Video Courses9 Jan 2023 0 repositories listed