Browse State-of-the-Art › Speech Synthesis › Papers, page 5
Speech Synthesis
Papers archive 2025-07-28
archive papers tagged: 1,249 · with a code link: 366 · where Syntology ran a sample: 101 (85 with a run with no instrument failure, 16 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (101 of 1,249 tagged: 85 with a run with no instrument failure, 16 where every run was a failure of Syntology's instrument)
Page 5 of 13: papers 401 to 500 of 1,249, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation29 Apr 2025 0 repositories listed
-
Towards Flow-Matching-based TTS without Classifier-Free Guidance29 Apr 2025 0 repositories listed
-
Generative Adversarial Network based Voice Conversion: Techniques, Challenges, and Recent Advancements27 Apr 2025 0 repositories listed
-
A Multi-Agent Framework for Automated Qinqiang Opera Script Generation Using Large Language Models22 Apr 2025 0 repositories listed
-
FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning22 Apr 2025 0 repositories listed
-
SOLIDO: A Robust Watermarking Method for Speech Synthesis via Low-Rank Adaptation21 Apr 2025 0 repositories listed
-
Collective Learning Mechanism based Optimal Transport Generative Adversarial Network for Non-parallel Voice Conversion18 Apr 2025 0 repositories listed
-
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis14 Apr 2025 0 repositories listed
-
Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis14 Apr 2025 0 repositories listed
-
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis12 Apr 2025 0 repositories listed
-
Empowering Global Voices: A Data-Efficient, Phoneme-Tone Adaptive Approach to High-Fidelity Speech Synthesis10 Apr 2025 0 repositories listed
-
SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow10 Apr 2025 0 repositories listed
-
VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models3 Apr 2025 0 repositories listed
-
SupertonicTTS: Towards Highly Scalable and Efficient Text-to-Speech System29 Mar 2025 0 repositories listed
-
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech21 Mar 2025 0 repositories listed
-
Good practices for evaluation of synthesized speech5 Mar 2025 0 repositories listed
-
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology3 Mar 2025 0 repositories listed
-
DiffCSS: Diverse and Expressive Conversational Speech Synthesis with Diffusion Models27 Feb 2025 0 repositories listed
-
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis26 Feb 2025 0 repositories listed
-
Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM24 Feb 2025 0 repositories listed
-
AV-Flow: Transforming Text to Audio-Visual Human-like Interactions18 Feb 2025 0 repositories listed
-
High-Fidelity Music Vocoder using Neural Audio Codecs18 Feb 2025 0 repositories listed
-
A Survey on Bridging EEG Signals and Generative AI: From Image and Text to Beyond17 Feb 2025 0 repositories listed
-
NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing17 Feb 2025 0 repositories listed
-
FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching16 Feb 2025 0 repositories listed
-
ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech13 Feb 2025 0 repositories listed
-
LoRP-TTS: Low-Rank Personalized Text-To-Speech11 Feb 2025 0 repositories listed
-
Non-invasive electromyographic speech neuroprosthesis: a geometric perspective9 Feb 2025 0 repositories listed
-
Gender Bias in Instruction-Guided Speech Synthesis Models8 Feb 2025 0 repositories listed
-
Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis3 Feb 2025 0 repositories listed
-
Compact Neural TTS Voices for Accessibility28 Jan 2025 0 repositories listed
-
Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation24 Jan 2025 0 repositories listed
-
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement23 Jan 2025 0 repositories listed
-
A Non-autoregressive Model for Joint STT and TTS15 Jan 2025 0 repositories listed
-
Speech Synthesis along Perceptual Voice Quality Dimensions15 Jan 2025 0 repositories listed
-
Exploring the encoding of linguistic representations in the Fully-Connected Layer of generative CNNs for Speech13 Jan 2025 0 repositories listed
-
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron10 Jan 2025 0 repositories listed
-
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control10 Jan 2025 0 repositories listed
-
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer10 Jan 2025 0 repositories listed
-
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis9 Jan 2025 0 repositories listed
-
Probing Speaker-specific Features in Speaker Representations9 Jan 2025 0 repositories listed
-
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts8 Jan 2025 0 repositories listed
-
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation28 Dec 2024 0 repositories listed
-
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting28 Dec 2024 0 repositories listed
-
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis26 Dec 2024 0 repositories listed
-
Intra- and Inter-modal Context Interaction Modeling for Conversational Speech Synthesis25 Dec 2024 0 repositories listed
-
MRI2Speech: Speech Synthesis from Articulatory Movements Recorded by Real-time MRI25 Dec 2024 0 repositories listed
-
Autoregressive Speech Synthesis with Next-Distribution Prediction22 Dec 2024 0 repositories listed
-
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis22 Dec 2024 0 repositories listed
-
Deep Speech Synthesis from Multimodal Articulatory Representations17 Dec 2024 0 repositories listed
-
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis16 Dec 2024 0 repositories listed
-
AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation13 Dec 2024 0 repositories listed
-
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens13 Dec 2024 0 repositories listed
-
Zero-Shot Mono-to-Binaural Speech Synthesis11 Dec 2024 0 repositories listed
-
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model4 Dec 2024 0 repositories listed
-
Visatronic: A Multimodal Decoder-Only Model for Speech Synthesis26 Nov 2024 0 repositories listed
-
VQalAttent: a Transparent Speech Generation Pipeline based on Transformer-learned VQ-VAE Latent Space22 Nov 2024 0 repositories listed
-
Debatts: Zero-Shot Debating Text-to-Speech Synthesis10 Nov 2024 0 repositories listed
-
Complete reconstruction of the tongue contour through acoustic to articulatory inversion using real-time MRI data4 Nov 2024 0 repositories listed
-
Augmenting Polish Automatic Speech Recognition System With Synthetic Data30 Oct 2024 0 repositories listed
-
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding29 Oct 2024 0 repositories listed
-
Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation27 Oct 2024 0 repositories listed
-
Making Social Platforms Accessible: Emotion-Aware Speech Generation with Integrated Text Analysis24 Oct 2024 0 repositories listed
-
Continuous Speech Synthesis using per-token Latent Diffusion21 Oct 2024 0 repositories listed
-
A Unified Framework for Collecting Text-to-Speech Synthesis Datasets for 22 Indian Languages18 Oct 2024 0 repositories listed
-
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding17 Oct 2024 0 repositories listed
-
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech17 Oct 2024 0 repositories listed
-
DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis17 Oct 2024 0 repositories listed
-
Beyond Oversmoothing: Evaluating DDPM and MSE for Scalable Speech Synthesis in ASR16 Oct 2024 0 repositories listed
-
DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis14 Oct 2024 0 repositories listed
-
Everyday Speech in the Indian Subcontinent14 Oct 2024 0 repositories listed
-
Bahasa Harmony: A Comprehensive Dataset for Bahasa Text-to-Speech Synthesis with Discrete Codec Modeling of EnGen-TTS9 Oct 2024 0 repositories listed
-
Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch9 Oct 2024 0 repositories listed
-
HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis6 Oct 2024 0 repositories listed
-
Adversarial Attacks and Robust Defenses in Speaker Embedding based Zero-Shot Text-to-Speech System5 Oct 2024 0 repositories listed
-
Generative Semantic Communication for Text-to-Speech Synthesis4 Oct 2024 0 repositories listed
-
MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech4 Oct 2024 0 repositories listed
-
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS30 Sep 2024 0 repositories listed
-
Quantitative Analysis of Audio-Visual Tasks: An Information-Theoretic Perspective29 Sep 2024 0 repositories listed
-
EmoPro: A Prompt Selection Strategy for Emotional Expression in LM-based Speech Synthesis27 Sep 2024 0 repositories listed
-
Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech24 Sep 2024 0 repositories listed
-
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis24 Sep 2024 0 repositories listed
-
NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization19 Sep 2024 0 repositories listed
-
Enhancing Multilingual Speech Generation and Recognition Abilities in LLMs with Constructed Code-switched Data17 Sep 2024 0 repositories listed
-
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation17 Sep 2024 0 repositories listed
-
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization16 Sep 2024 0 repositories listed
-
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion16 Sep 2024 0 repositories listed
-
Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation14 Sep 2024 0 repositories listed
-
LLM-Powered Grapheme-to-Phoneme Conversion: Benchmark and Case Study13 Sep 2024 0 repositories listed
-
Text-To-Speech Synthesis In The Wild13 Sep 2024 0 repositories listed
-
Full-text Error Correction for Chinese Speech Recognition with Large Language Model12 Sep 2024 0 repositories listed
-
Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach10 Sep 2024 0 repositories listed
-
What happens to diffusion model likelihood when your model is conditional?10 Sep 2024 0 repositories listed
-
AS-Speech: Adaptive Style For Speech Synthesis9 Sep 2024 0 repositories listed
-
Fast, High-Quality and Parameter-Efficient Articulatory Synthesis using Differentiable DSP4 Sep 2024 0 repositories listed
-
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders3 Sep 2024 0 repositories listed
-
VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka3 Sep 2024 0 repositories listed
-
SelectTTS: Synthesizing Anyone's Voice via Discrete Unit-Based Frame Selection30 Aug 2024 0 repositories listed
-
Literary and Colloquial Dialect Identification for Tamil using Acoustic Features27 Aug 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.