Browse State-of-the-Art › Text to Speech › Papers, page 6
Text to Speech
Papers archive 2025-07-28
archive papers tagged: 1,419 · with a code link: 399 · where Syntology ran a sample: 108 (96 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (108 of 1,419 tagged: 96 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument)
Page 6 of 15: papers 501 to 600 of 1,419, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
LoRP-TTS: Low-Rank Personalized Text-To-Speech11 Feb 2025 0 repositories listed
-
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement11 Feb 2025 0 repositories listed
-
Speech to Speech Translation with Translatotron: A State of the Art Review9 Feb 2025 0 repositories listed
-
Gender Bias in Instruction-Guided Speech Synthesis Models8 Feb 2025 0 repositories listed
-
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech5 Feb 2025 0 repositories listed
-
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation4 Feb 2025 0 repositories listed
-
EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis2 Feb 2025 0 repositories listed
-
VisualSpeech: Enhance Prosody with Visual Context in TTS31 Jan 2025 0 repositories listed
-
BreezyVoice: Adapting TTS for Taiwanese Mandarin with Enhanced Polyphone Disambiguation -- Challenges and Insights29 Jan 2025 0 repositories listed
-
Compact Neural TTS Voices for Accessibility28 Jan 2025 0 repositories listed
-
Characteristic-Specific Partial Fine-Tuning for Efficient Emotion and Speaker Adaptation in Codec Language Text-to-Speech Models24 Jan 2025 0 repositories listed
-
Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation24 Jan 2025 0 repositories listed
-
LoCoML: A Framework for Real-World ML Inference Pipelines24 Jan 2025 0 repositories listed
-
Generative Data Augmentation Challenge: Zero-Shot Speech Synthesis for Personalized Speech Enhancement23 Jan 2025 0 repositories listed
-
Development of an Inclusive Educational Platform Using Open Technologies and Machine Learning: A Case Study on Accessibility Enhancement22 Jan 2025 0 repositories listed
-
A Domain Adaptation Framework for Speech Recognition Systems with Only Synthetic data21 Jan 2025 0 repositories listed
-
Speech Synthesis along Perceptual Voice Quality Dimensions15 Jan 2025 0 repositories listed
-
Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement15 Jan 2025 0 repositories listed
-
AI-Powered Assistive Technologies for Visual Impairment14 Jan 2025 0 repositories listed
-
Low-Resource Text-to-Speech Synthesis Using Noise-Augmented Training of ForwardTacotron10 Jan 2025 0 repositories listed
-
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model10 Jan 2025 0 repositories listed
-
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction10 Jan 2025 0 repositories listed
-
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control10 Jan 2025 0 repositories listed
-
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer10 Jan 2025 0 repositories listed
-
Probing Speaker-specific Features in Speaker Representations9 Jan 2025 0 repositories listed
-
Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model8 Jan 2025 0 repositories listed
-
Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT2 Jan 2025 0 repositories listed
-
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting28 Dec 2024 0 repositories listed
-
"I've Heard of You!": Generate Spoken Named Entity Recognition Data for Unseen Entities26 Dec 2024 0 repositories listed
-
Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID26 Dec 2024 0 repositories listed
-
Advancing NAM-to-Speech Conversion with Novel Methods and the MultiNAM Dataset25 Dec 2024 0 repositories listed
-
Autoregressive Speech Synthesis with Next-Distribution Prediction22 Dec 2024 0 repositories listed
-
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis22 Dec 2024 0 repositories listed
-
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective22 Dec 2024 0 repositories listed
-
Interleaved Speech-Text Language Models are Simple Streaming Text to Speech Synthesizers20 Dec 2024 0 repositories listed
-
Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling19 Dec 2024 0 repositories listed
-
Enhancing Naturalness in LLM-Generated Utterances through Disfluency Insertion17 Dec 2024 0 repositories listed
-
Phoneme-Level Feature Discrepancies: A Key to Detecting Sophisticated Speech Deepfakes17 Dec 2024 0 repositories listed
-
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis16 Dec 2024 0 repositories listed
-
AMuSeD: An Attentive Deep Neural Network for Multimodal Sarcasm Detection Incorporating Bi-modal Data Augmentation13 Dec 2024 0 repositories listed
-
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens13 Dec 2024 0 repositories listed
-
CSSinger: End-to-End Chunkwise Streaming Singing Voice Synthesis System Based on Conditional Variational Autoencoder12 Dec 2024 0 repositories listed
-
A Preliminary Analysis of Automatic Word and Syllable Prominence Detection in Non-Native Speech With Text-to-Speech Prosody Embeddings11 Dec 2024 0 repositories listed
-
A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction11 Dec 2024 0 repositories listed
-
Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration11 Dec 2024 0 repositories listed
-
LatentSpeech: Latent Diffusion for Text-To-Speech Generation11 Dec 2024 0 repositories listed
-
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations9 Dec 2024 0 repositories listed
-
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles4 Dec 2024 0 repositories listed
-
Text Is Not All You Need: Multimodal Prompting Helps LLMs Understand Humor1 Dec 2024 0 repositories listed
-
Continual Learning in Machine Speech Chain Using Gradient Episodic Memory27 Nov 2024 0 repositories listed
-
SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation27 Nov 2024 0 repositories listed
-
Visatronic: A Multimodal Decoder-Only Model for Speech Synthesis26 Nov 2024 0 repositories listed
-
Hard-Synth: Synthesizing Diverse Hard Samples for ASR using Zero-Shot TTS and LLM20 Nov 2024 0 repositories listed
-
A Context-Based Numerical Format Prediction for a Text-To-Speech System19 Nov 2024 0 repositories listed
-
Leveraging Virtual Reality and AI Tutoring for Language Learning: A Case Study of a Virtual Campus Environment with OpenAI GPT Integration with Unity 3D19 Nov 2024 0 repositories listed
-
Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation19 Nov 2024 0 repositories listed
-
Improving Grapheme-to-Phoneme Conversion through In-Context Knowledge Retrieval with Large Language Models12 Nov 2024 0 repositories listed
-
Debatts: Zero-Shot Debating Text-to-Speech Synthesis10 Nov 2024 0 repositories listed
-
CUIfy the XR: An Open-Source Package to Embed LLM-powered Conversational Agents in XR7 Nov 2024 0 repositories listed
-
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?31 Oct 2024 0 repositories listed
-
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding29 Oct 2024 0 repositories listed
-
RDSinger: Reference-based Diffusion Network for Singing Voice Synthesis29 Oct 2024 0 repositories listed
-
Asynchronous Tool Usage for Real-Time Agents28 Oct 2024 0 repositories listed
-
Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation27 Oct 2024 0 repositories listed
-
Evaluating and Improving Automatic Speech Recognition Systems for Korean Meteorological Experts24 Oct 2024 0 repositories listed
-
Making Social Platforms Accessible: Emotion-Aware Speech Generation with Integrated Text Analysis24 Oct 2024 0 repositories listed
-
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams23 Oct 2024 0 repositories listed
-
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap22 Oct 2024 0 repositories listed
-
Continuous Speech Synthesis using per-token Latent Diffusion21 Oct 2024 0 repositories listed
-
A Unified Framework for Collecting Text-to-Speech Synthesis Datasets for 22 Indian Languages18 Oct 2024 0 repositories listed
-
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech17 Oct 2024 0 repositories listed
-
DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis17 Oct 2024 0 repositories listed
-
Enhancing Crowdsourced Audio for Text-to-Speech Models17 Oct 2024 0 repositories listed
-
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation17 Oct 2024 0 repositories listed
-
ERVQ: Enhanced Residual Vector Quantization with Intra-and-Inter-Codebook Optimization for Neural Audio Codecs16 Oct 2024 0 repositories listed
-
DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis14 Oct 2024 0 repositories listed
-
IsoChronoMeter: A simple and effective isochronic translation evaluation metric14 Oct 2024 0 repositories listed
-
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling12 Oct 2024 0 repositories listed
-
Unsupervised Data Validation Methods for Efficient Model Training10 Oct 2024 0 repositories listed
-
Bahasa Harmony: A Comprehensive Dataset for Bahasa Text-to-Speech Synthesis with Discrete Codec Modeling of EnGen-TTS9 Oct 2024 0 repositories listed
-
Can DeepFake Speech be Reliably Detected?9 Oct 2024 0 repositories listed
-
Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch9 Oct 2024 0 repositories listed
-
SegINR: Segment-wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech7 Oct 2024 0 repositories listed
-
HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis6 Oct 2024 0 repositories listed
-
Adversarial Attacks and Robust Defenses in Speaker Embedding based Zero-Shot Text-to-Speech System5 Oct 2024 0 repositories listed
-
Generative Semantic Communication for Text-to-Speech Synthesis4 Oct 2024 0 repositories listed
-
MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech4 Oct 2024 0 repositories listed
-
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens4 Oct 2024 0 repositories listed
-
Augmentation through Laundering Attacks for Audio Spoof Detection1 Oct 2024 0 repositories listed
-
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS30 Sep 2024 0 repositories listed
-
Word-wise intonation model for cross-language TTS systems30 Sep 2024 0 repositories listed
-
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control26 Sep 2024 0 repositories listed
-
Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions25 Sep 2024 0 repositories listed
-
Exploring synthetic data for cross-speaker style transfer in style representation based TTS25 Sep 2024 0 repositories listed
-
Beyond Text-to-Text: An Overview of Multimodal and Generative Artificial Intelligence for Education Using Topic Modeling24 Sep 2024 0 repositories listed
-
Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech24 Sep 2024 0 repositories listed
-
StyleFusion TTS: Multimodal Style-control and Enhanced Feature Fusion for Zero-shot Text-to-speech Synthesis24 Sep 2024 0 repositories listed
-
On the Feasibility of Fully AI-automated Vishing Attacks20 Sep 2024 0 repositories listed
-
Zero-shot Cross-lingual Voice Transfer for TTS20 Sep 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.