Browse State-of-the-Art › Speech Synthesis › Papers, page 7
Speech Synthesis
Papers archive 2025-07-28
archive papers tagged: 1,249 · with a code link: 366 · where Syntology ran a sample: 101 (85 with a run with no instrument failure, 16 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (101 of 1,249 tagged: 85 with a run with no instrument failure, 16 where every run was a failure of Syntology's instrument)
Page 7 of 13: papers 601 to 700 of 1,249, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
A Discourse-level Multi-scale Prosodic Model for Fine-grained Emotion Analysis21 Sep 2023 0 repositories listed
-
Speak While You Think: Streaming Speech Synthesis During Text Generation20 Sep 2023 0 repositories listed
-
Exploring Speech Enhancement for Low-resource Speech Synthesis19 Sep 2023 0 repositories listed
-
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition19 Sep 2023 0 repositories listed
-
Corpus Synthesis for Zero-shot ASR domain Adaptation using Large Language Models18 Sep 2023 0 repositories listed
-
Speech Synthesis By Unrolling Diffusion Process using Neural Network Layers18 Sep 2023 0 repositories listed
-
Cross-lingual Knowledge Distillation via Flow-based Voice Conversion for Robust Polyglot Text-To-Speech15 Sep 2023 0 repositories listed
-
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks14 Sep 2023 0 repositories listed
-
12 Sep 2023 0 repositories listed
-
Cross-Utterance Conditioned VAE for Speech Generation8 Sep 2023 0 repositories listed
-
MuLanTTS: The Microsoft Speech Synthesis System for Blizzard Challenge 20236 Sep 2023 0 repositories listed
-
The FruitShell French synthesis system at the Blizzard 2023 Challenge1 Sep 2023 0 repositories listed
-
Towards Spontaneous Style Modeling with Semi-supervised Pre-training for Conversational Text-to-Speech Synthesis31 Aug 2023 0 repositories listed
-
The DeepZen Speech Synthesis System for Blizzard Challenge 202330 Aug 2023 0 repositories listed
-
Generalizable Zero-Shot Speaker Adaptive Speech Synthesis with Disentangled Representations24 Aug 2023 0 repositories listed
-
TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition21 Aug 2023 0 repositories listed
-
Accurate synthesis of Dysarthric Speech for ASR data augmentation16 Aug 2023 0 repositories listed
-
AffectEcho: Speaker Independent and Language-Agnostic Emotion and Affect Transfer for Speech Synthesis16 Aug 2023 0 repositories listed
-
iSTFTNet2: Faster and More Lightweight iSTFT-Based Neural Vocoder Using 1D-2D CNN14 Aug 2023 0 repositories listed
-
EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis10 Aug 2023 0 repositories listed
-
On Error Propagation of Diffusion Models9 Aug 2023 0 repositories listed
-
SALTTS: Leveraging Self-Supervised Speech Representations for improved Text-to-Speech Synthesis2 Aug 2023 0 repositories listed
-
Audio-visual video-to-speech synthesis with synthesized input audio31 Jul 2023 0 repositories listed
-
Comparing normalizing flows and diffusion models for prosody and acoustic modelling in text-to-speech31 Jul 2023 0 repositories listed
-
METTS: Multilingual Emotional Text-to-Speech by Cross-speaker and Cross-lingual Emotion Transfer29 Jul 2023 0 repositories listed
-
Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic Coding28 Jul 2023 0 repositories listed
-
An analysis on the effects of speaker embedding choice in non auto-regressive TTS19 Jul 2023 0 repositories listed
-
SLMGAN: Exploiting Speech Language Model Representations for Unsupervised Zero-Shot Voice Conversion in GANs18 Jul 2023 0 repositories listed
-
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis14 Jul 2023 0 repositories listed
-
On the Use of Self-Supervised Speech Representations in Spontaneous Speech Synthesis11 Jul 2023 0 repositories listed
-
RobustL2S: Speaker-Specific Lip-to-Speech Synthesis exploiting Self-Supervised Representations3 Jul 2023 0 repositories listed
-
High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units29 Jun 2023 0 repositories listed
-
Large-scale unsupervised audio pre-training for video-to-speech synthesis27 Jun 2023 0 repositories listed
-
DSE-TTS: Dual Speaker Embedding for Cross-Lingual Text-to-Speech25 Jun 2023 0 repositories listed
-
Strategies in Transfer Learning for Low-Resource Speech Synthesis: Phone Mapping, Features Input, and Source Language Selection21 Jun 2023 0 repositories listed
-
Visual-Aware Text-to-Speech21 Jun 2023 0 repositories listed
-
Cross-lingual Prosody Transfer for Expressive Machine Dubbing20 Jun 2023 0 repositories listed
-
CML-TTS A Multilingual Dataset for Speech Synthesis in Low-Resource Languages16 Jun 2023 0 repositories listed
-
Investigating the Utility of Surprisal from Large Language Models for Speech Synthesis Prosody16 Jun 2023 0 repositories listed
-
Diff-TTSG: Denoising probabilistic integrated speech and gesture synthesis15 Jun 2023 0 repositories listed
-
PauseSpeech: Natural Speech Synthesis via Pre-trained Language Model and Pause-based Prosody Modeling13 Jun 2023 0 repositories listed
-
HiddenSinger: High-Quality Singing Voice Synthesis via Neural Audio Codec and Latent Diffusion Models12 Jun 2023 0 repositories listed
-
Boosting Fast and High-Quality Speech Synthesis with Linear Diffusion9 Jun 2023 0 repositories listed
-
PolyVoice: Language Models for Speech to Speech Translation5 Jun 2023 0 repositories listed
-
Rhythm-controllable Attention with High Robustness for Long Sentence Speech Synthesis5 Jun 2023 0 repositories listed
-
Speaker-independent neural formant synthesis2 Jun 2023 0 repositories listed
-
Speech inpainting: Context-based speech synthesis guided by video1 Jun 2023 0 repositories listed
-
Text-to-Speech Pipeline for Swiss German -- A comparison31 May 2023 0 repositories listed
-
Automatic Evaluation of Turn-taking Cues in Conversational Speech Synthesis29 May 2023 0 repositories listed
-
Creating Personalized Synthetic Voices from Post-Glossectomy Speech with Guided Diffusion Models27 May 2023 0 repositories listed
-
Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM24 May 2023 0 repositories listed
-
CALLS: Japanese Empathetic Dialogue Speech Corpus of Complaint Handling and Attentive Listening in Customer Center23 May 2023 0 repositories listed
-
ChatGPT-EDSS: Empathetic Dialogue Speech Synthesis Trained from ChatGPT-derived Context Word Embeddings23 May 2023 0 repositories listed
-
ZET-Speech: Zero-shot adaptive Emotion-controllable Text-to-Speech Synthesis with Diffusion and Style-based Models23 May 2023 0 repositories listed
-
Text Generation with Speech Synthesis for ASR Data Augmentation22 May 2023 0 repositories listed
-
VAKTA-SETU: A Speech-to-Speech Machine Translation Service in Select Indic Languages21 May 2023 0 repositories listed
-
MParrotTTS: Multilingual Multi-speaker Text to Speech Synthesis in Low Resource Setting19 May 2023 0 repositories listed
-
A unified front-end framework for English text-to-speech synthesis18 May 2023 0 repositories listed
-
Empirical Analysis of Oral and Nasal Vowels of Konkani17 May 2023 0 repositories listed
-
Zero-shot personalized lip-to-speech synthesis with face image based voice control9 May 2023 0 repositories listed
-
Accented Text-to-Speech Synthesis with Limited Data8 May 2023 0 repositories listed
-
M2-CTTS: End-to-End Multi-scale Multi-modal Conversational Text-to-Speech Synthesis3 May 2023 0 repositories listed
-
A Review of Deep Learning Techniques for Speech Processing30 Apr 2023 0 repositories listed
-
Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model24 Apr 2023 0 repositories listed
-
Ensemble prosody prediction for expressive speech synthesis3 Apr 2023 0 repositories listed
-
Text is All You Need: Personalizing ASR Models using Controllable Speech Synthesis27 Mar 2023 0 repositories listed
-
Wave-U-Net Discriminator: Fast and Lightweight Discriminator for Generative Adversarial Network-Based Speech Synthesis24 Mar 2023 0 repositories listed
-
A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI23 Mar 2023 0 repositories listed
-
Transformers in Speech Processing: A Survey21 Mar 2023 0 repositories listed
-
Controllable Prosody Generation With Partial Inputs14 Mar 2023 0 repositories listed
-
Improving Prosody for Cross-Speaker Style Transfer by Semi-Supervised Style Extractor and Hierarchical Modeling in Speech Synthesis14 Mar 2023 0 repositories listed
-
QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis14 Mar 2023 0 repositories listed
-
VANI: Very-lightweight Accent-controllable TTS for Native and Non-native speakers with Identity Preservation14 Mar 2023 0 repositories listed
-
Do Prosody Transfer Models Transfer Prosody?7 Mar 2023 0 repositories listed
-
FoundationTTS: Text-to-Speech for ASR Customization with Generative Language Model6 Mar 2023 0 repositories listed
-
DTW-SiameseNet: Dynamic Time Warped Siamese Network for Mispronunciation Detection and Correction1 Mar 2023 0 repositories listed
-
On the Audio-visual Synchronization for Lip-to-Speech Synthesis1 Mar 2023 0 repositories listed
-
ParrotTTS: Text-to-Speech synthesis by exploiting self-supervised representations1 Mar 2023 0 repositories listed
-
ClArTTS: An Open-Source Classical Arabic Text-to-Speech Corpus28 Feb 2023 0 repositories listed
-
CrossSpeech: Speaker-independent Acoustic Representation for Cross-lingual Speech Synthesis28 Feb 2023 0 repositories listed
-
UniFLG: Unified Facial Landmark Generator from Text or Speech28 Feb 2023 0 repositories listed
-
Fast and small footprint Hybrid HMM-HiFiGAN based system for speech synthesis in Indian languages13 Feb 2023 0 repositories listed
-
Beyond Statistical Similarity: Rethinking Metrics for Deep Generative Models in Engineering Design6 Feb 2023 0 repositories listed
-
UzbekTagger: The rule-based POS tagger for Uzbek language30 Jan 2023 0 repositories listed
-
On granularity of prosodic representations in expressive text-to-speech26 Jan 2023 0 repositories listed
-
Multilingual Multiaccented Multispeaker TTS with RADTTS24 Jan 2023 0 repositories listed
-
Regeneration Learning: A Learning Paradigm for Data Generation21 Jan 2023 0 repositories listed
-
Applying Automated Machine Translation to Educational Video Courses9 Jan 2023 0 repositories listed
-
ReVISE: Self-Supervised Speech Resynthesis With Visual Input for Universal and Generalized Speech Regeneration1 Jan 2023 0 repositories listed
-
HMM-based data augmentation for E2E systems for building conversational speech synthesis systems22 Dec 2022 0 repositories listed
-
21 Dec 2022 0 repositories listed
-
Investigation of Japanese PnG BERT language model in text-to-speech synthesis for pitch accent language16 Dec 2022 0 repositories listed
-
Text-to-speech synthesis based on latent variable conversion using diffusion probabilistic model and variational autoencoder16 Dec 2022 0 repositories listed
-
Style-Label-Free: Cross-Speaker Style Transfer by Quantized VAE and Speaker-wise Normalization in Speech Synthesis13 Dec 2022 0 repositories listed
-
SNAC: Speaker-normalized affine coupling layer in flow-based architecture for zero-shot multi-speaker text-to-speech30 Nov 2022 0 repositories listed
-
Controllable speech synthesis by learning discrete phoneme-level prosodic representations29 Nov 2022 0 repositories listed
-
Contextual Expressive Text-to-Speech26 Nov 2022 0 repositories listed
-
Efficient Incremental Text-to-Speech on GPUs25 Nov 2022 0 repositories listed
-
LA-VocE: Low-SNR Audio-visual Speech Enhancement using Neural Vocoders20 Nov 2022 0 repositories listed
-
Multi-Speaker Expressive Speech Synthesis via Multiple Factors Decoupling19 Nov 2022 0 repositories listed