Browse State-of-the-Art › Text to Speech › Papers, page 5
Text to Speech
Papers archive 2025-07-28
archive papers tagged: 1,419 · with a code link: 399 · where Syntology ran a sample: 108 (96 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (108 of 1,419 tagged: 96 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument)
Page 5 of 15: papers 401 to 500 of 1,419, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech17 Jul 2025 0 repositories listed
-
P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge15 Jul 2025 0 repositories listed
-
An Empirical Evaluation of AI-Powered Non-Player Characters' Perceived Realism and Performance in Virtual Reality Environments14 Jul 2025 0 repositories listed
-
Exploiting Leaderboards for Large-Scale Distribution of Malicious Models11 Jul 2025 0 repositories listed
-
MIDI-VALLE: Improving Expressive Piano Performance Synthesis Through Neural Codec Language Modelling11 Jul 2025 0 repositories listed
-
Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis8 Jul 2025 0 repositories listed
-
An Exploration of ECAPA-TDNN and x-vector Speaker Representations in Zero-shot Multi-speaker TTS25 Jun 2025 0 repositories listed
-
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems24 Jun 2025 0 repositories listed
-
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization20 Jun 2025 0 repositories listed
-
Optimizing Multilingual Text-To-Speech with Accents & Emotions19 Jun 2025 0 repositories listed
-
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement19 Jun 2025 0 repositories listed
-
PredGen: Accelerated Inference of Large Language Models through Input-Time Speculation for Real-Time Speech Interaction18 Jun 2025 0 repositories listed
-
Phonikud: Hebrew Grapheme-to-Phoneme Conversion for Real-Time Text-to-Speech14 Jun 2025 0 repositories listed
-
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling14 Jun 2025 0 repositories listed
-
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs12 Jun 2025 0 repositories listed
-
S2ST-Omni: An Efficient and Scalable Multilingual Speech-to-Speech Translation Framework via Seamless Speech-Text Alignment and Streaming Speech Generation11 Jun 2025 0 repositories listed
-
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching11 Jun 2025 0 repositories listed
-
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data10 Jun 2025 0 repositories listed
-
Seeing Voices: Generating A-Roll Video from Audio with Mirage9 Jun 2025 0 repositories listed
-
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation9 Jun 2025 0 repositories listed
-
Voice Impression Control in Zero-Shot TTS6 Jun 2025 0 repositories listed
-
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning5 Jun 2025 0 repositories listed
-
Intelligibility of Text-to-Speech Systems for Mathematical Expressions5 Jun 2025 0 repositories listed
-
A Novel Data Augmentation Approach for Automatic Speaking Assessment on Opinion Expressions4 Jun 2025 0 repositories listed
-
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing4 Jun 2025 0 repositories listed
-
Can we reconstruct a dysarthric voice with the large speech model Parler TTS?4 Jun 2025 0 repositories listed
-
HiFiTTS-2: A Large-Scale High Bandwidth Speech Dataset4 Jun 2025 0 repositories listed
-
UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation4 Jun 2025 0 repositories listed
-
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech3 Jun 2025 0 repositories listed
-
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions3 Jun 2025 0 repositories listed
-
Towards a Japanese Full-duplex Spoken Dialogue System3 Jun 2025 0 repositories listed
-
SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction2 Jun 2025 0 repositories listed
-
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing2 Jun 2025 0 repositories listed
-
Zero-Shot Text-to-Speech for Vietnamese2 Jun 2025 0 repositories listed
-
Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models1 Jun 2025 0 repositories listed
-
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems31 May 2025 0 repositories listed
-
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation30 May 2025 0 repositories listed
-
Werewolf: A Straightforward Game Framework with TTS for Improved User Engagement30 May 2025 0 repositories listed
-
Can Emotion Fool Anti-spoofing?29 May 2025 0 repositories listed
-
LLM-Synth4KWS: Scalable Automatic Generation and Synthesis of Confusable Data for Custom Keyword Spotting29 May 2025 0 repositories listed
-
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech27 May 2025 0 repositories listed
-
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling26 May 2025 0 repositories listed
-
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech26 May 2025 0 repositories listed
-
KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization26 May 2025 0 repositories listed
-
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling26 May 2025 0 repositories listed
-
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning25 May 2025 0 repositories listed
-
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis25 May 2025 0 repositories listed
-
SpeakStream: Streaming Text-to-Speech with Interleaved Data25 May 2025 0 repositories listed
-
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt24 May 2025 0 repositories listed
-
RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations24 May 2025 0 repositories listed
-
What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection23 May 2025 0 repositories listed
-
Benchmarking Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS222 May 2025 0 repositories listed
-
MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling21 May 2025 0 repositories listed
-
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information21 May 2025 0 repositories listed
-
Voicing Personas: Rewriting Persona Descriptions into Style Prompts for Controllable Text-to-Speech21 May 2025 0 repositories listed
-
AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models20 May 2025 0 repositories listed
-
FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation20 May 2025 0 repositories listed
-
Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English20 May 2025 0 repositories listed
-
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising20 May 2025 0 repositories listed
-
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement20 May 2025 0 repositories listed
-
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching19 May 2025 0 repositories listed
-
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis18 May 2025 0 repositories listed
-
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese16 May 2025 0 repositories listed
-
UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech15 May 2025 0 repositories listed
-
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications12 May 2025 0 repositories listed
-
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder12 May 2025 0 repositories listed
-
Bridging the Gap: An Intermediate Language for Enhanced and Cost-Effective Grapheme-to-Phoneme Conversion with Homographs with Multiple Pronunciations Disambiguation10 May 2025 0 repositories listed
-
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech8 May 2025 0 repositories listed
-
Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations8 May 2025 0 repositories listed
-
Generating Narrated Lecture Videos from Slides with Synchronized Highlights5 May 2025 0 repositories listed
-
Sadeed: Advancing Arabic Diacritization Through Small Language Model30 Apr 2025 0 repositories listed
-
Towards Flow-Matching-based TTS without Classifier-Free Guidance29 Apr 2025 0 repositories listed
-
A Multi-Agent Framework for Automated Qinqiang Opera Script Generation Using Large Language Models22 Apr 2025 0 repositories listed
-
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting17 Apr 2025 0 repositories listed
-
GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM15 Apr 2025 0 repositories listed
-
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis14 Apr 2025 0 repositories listed
-
Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis14 Apr 2025 0 repositories listed
-
Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation11 Apr 2025 0 repositories listed
-
Empowering Global Voices: A Data-Efficient, Phoneme-Tone Adaptive Approach to High-Fidelity Speech Synthesis10 Apr 2025 0 repositories listed
-
SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow10 Apr 2025 0 repositories listed
-
SpeakEasy: Enhancing Text-to-Speech Interactions for Expressive Content Creation7 Apr 2025 0 repositories listed
-
Speculative End-Turn Detector for Efficient Speech Chatbot Assistant30 Mar 2025 0 repositories listed
-
SupertonicTTS: Towards Highly Scalable and Efficient Text-to-Speech System29 Mar 2025 0 repositories listed
-
DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation28 Mar 2025 0 repositories listed
-
Dual Audio-Centric Modality Coupling for Talking Head Generation26 Mar 2025 0 repositories listed
-
Your voice is your voice: Supporting Self-expression through Speech Generation and LLMs in Augmented and Alternative Communication21 Mar 2025 0 repositories listed
-
MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation14 Mar 2025 0 repositories listed
-
An Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASR11 Mar 2025 0 repositories listed
-
VocalEyes: Enhancing Environmental Perception for the Visually Impaired through Vision-Language Models and Distance-Aware Object Detection10 Mar 2025 0 repositories listed
-
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training4 Mar 2025 0 repositories listed
-
Direct Speech to Speech Translation: A Review3 Mar 2025 0 repositories listed
-
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation2 Mar 2025 0 repositories listed
-
Telephone Surveys Meet Conversational AI: Evaluating a LLM-Based Telephone Survey System at Scale27 Feb 2025 0 repositories listed
-
Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding26 Feb 2025 0 repositories listed
-
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision26 Feb 2025 0 repositories listed
-
MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis26 Feb 2025 0 repositories listed
-
Balancing Speech Understanding and Generation Using Continual Pre-training for Codec-based Speech LLM24 Feb 2025 0 repositories listed
-
NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing17 Feb 2025 0 repositories listed
-
SyncSpeech: Low-Latency and Efficient Dual-Stream Text-to-Speech based on Temporal Masked Transformer16 Feb 2025 0 repositories listed
-
ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech13 Feb 2025 0 repositories listed