Browse State-of-the-Art › text-to-speech › Papers, page 8
text-to-speech
Papers archive 2025-07-28
archive papers tagged: 1,413 · with a code link: 395 · where Syntology ran a sample: 106 (95 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (106 of 1,413 tagged: 95 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument)
Page 8 of 15: papers 701 to 800 of 1,413, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Multi-speaker Text-to-speech Training with Speaker Anonymized Data20 May 2024 0 repositories listed
-
VR-GPT: Visual Language Model for Intelligent Virtual Reality Applications19 May 2024 0 repositories listed
-
Exploring speech style spaces with language models: Emotional TTS without emotion labels18 May 2024 0 repositories listed
-
Building a Luganda Text-to-Speech Model From Crowdsourced Data16 May 2024 0 repositories listed
-
Evaluating Text-to-Speech Synthesis from a Large Discrete Token-based Speech Language Model16 May 2024 0 repositories listed
-
Faces that Speak: Jointly Synthesising Talking Face and Speech from Text16 May 2024 0 repositories listed
-
Towards Evaluating the Robustness of Automatic Speech Recognition Systems via Audio Style Transfer15 May 2024 0 repositories listed
-
Real-Time Pill Identification for the Visually Impaired Using Deep Learning8 May 2024 0 repositories listed
-
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech30 Apr 2024 0 repositories listed
-
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality27 Apr 2024 0 repositories listed
-
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations23 Apr 2024 0 repositories listed
-
Retrieval-Augmented Audio Deepfake Detection22 Apr 2024 0 repositories listed
-
Prior-agnostic Multi-scale Contrastive Text-Audio Pre-training for Parallelized TTS Frontend Modeling14 Apr 2024 0 repositories listed
-
Voice-Assisted Real-Time Traffic Sign Recognition System Using Convolutional Neural Network11 Apr 2024 0 repositories listed
-
The X-LANCE Technical Report for Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge9 Apr 2024 0 repositories listed
-
Cross-Domain Audio Deepfake Detection: Dataset and Analysis7 Apr 2024 0 repositories listed
-
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis4 Apr 2024 0 repositories listed
-
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech3 Apr 2024 0 repositories listed
-
A Review of Multi-Modal Large Language and Vision Models28 Mar 2024 0 repositories listed
-
Isometric Neural Machine Translation using Phoneme Count Ratio Reward-based Reinforcement Learning20 Mar 2024 0 repositories listed
-
Creating an African American-Sounding TTS: Guidelines, Technical Challenges,and Surprising Evaluations17 Mar 2024 0 repositories listed
-
EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech13 Mar 2024 0 repositories listed
-
Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation7 Mar 2024 0 repositories listed
-
AttentionStitch: How Attention Solves the Speech Editing Problem5 Mar 2024 0 repositories listed
-
Towards Accurate Lip-to-Speech Synthesis in-the-Wild2 Mar 2024 0 repositories listed
-
Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data29 Feb 2024 0 repositories listed
-
Daisy-TTS: Simulating Wider Spectrum of Emotions via Prosody Embedding Decomposition22 Feb 2024 0 repositories listed
-
Efficient data selection employing Semantic Similarity-based Graph Structures for model training22 Feb 2024 0 repositories listed
-
On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models19 Feb 2024 0 repositories listed
-
Ain't Misbehavin' -- Using LLMs to Generate Expressive Robot Behavior in Conversations with the Tabletop Robot Haru18 Feb 2024 0 repositories listed
-
BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data12 Feb 2024 0 repositories listed
-
Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like12 Feb 2024 0 repositories listed
-
A New Approach to Voice Authenticity9 Feb 2024 0 repositories listed
-
Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations5 Feb 2024 0 repositories listed
-
Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech1 Feb 2024 0 repositories listed
-
MunTTS: A Text-to-Speech System for Mundari28 Jan 2024 0 repositories listed
-
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech25 Jan 2024 0 repositories listed
-
Maximizing Data Efficiency for Cross-Lingual TTS Adaptation by Self-Supervised Representation Mixing and Embedding Initialization23 Jan 2024 0 repositories listed
-
Empowering Communication: Speech Technology for Indian and Western Accents through AI-powered Speech Synthesis22 Jan 2024 0 repositories listed
-
Adversarial speech for voice privacy protection from Personalized Speech generation22 Jan 2024 0 repositories listed
-
Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech19 Jan 2024 0 repositories listed
-
MCMChaos: Improvising Rap Music with MCMC Methods and Chaos Theory15 Jan 2024 0 repositories listed
-
ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering14 Jan 2024 0 repositories listed
-
End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec211 Jan 2024 0 repositories listed
-
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters10 Jan 2024 0 repositories listed
-
Evaluating and Personalizing User-Perceived Quality of Text-to-Speech Voices for Delivering Mindfulness Meditation with Different Physical Embodiments7 Jan 2024 0 repositories listed
-
Transfer the linguistic representations from TTS to accent conversion with non-parallel data7 Jan 2024 0 repositories listed
-
Incremental FastPitch: Chunk-based High Quality Text to Speech3 Jan 2024 0 repositories listed
-
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction3 Jan 2024 0 repositories listed
-
Boosting Large Language Model for Speech Synthesis: An Empirical Study30 Dec 2023 0 repositories listed
-
Normalization of Lithuanian Text Using Regular Expressions29 Dec 2023 0 repositories listed
-
AE-Flow: AutoEncoder Normalizing Flow27 Dec 2023 0 repositories listed
-
Creating New Voices using Normalizing Flows22 Dec 2023 0 repositories listed
-
External Knowledge Augmented Polyphone Disambiguation Using Large Language Model19 Dec 2023 0 repositories listed
-
A review-based study on different Text-to-Speech technologies17 Dec 2023 0 repositories listed
-
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis17 Dec 2023 0 repositories listed
-
An Experimental Study: Assessing the Combined Framework of WavLM and BEST-RQ for Text-to-Speech Synthesis8 Dec 2023 0 repositories listed
-
Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis6 Dec 2023 0 repositories listed
-
Code-Mixed Text to Speech Synthesis under Low-Resource Constraints2 Dec 2023 0 repositories listed
-
Rapid Speaker Adaptation in Low Resource Text to Speech Systems using Synthetic Data and Transfer learning2 Dec 2023 0 repositories listed
-
Vulnerability of Automatic Identity Recognition to Audio-Visual Deepfakes29 Nov 2023 0 repositories listed
-
Guided Flows for Generative Modeling and Decision Making22 Nov 2023 0 repositories listed
-
Data Center Audio/Video Intelligence on Device (DAVID) -- An Edge-AI Platform for Smart-Toys18 Nov 2023 0 repositories listed
-
Utilizing Speech Emotion Recognition and Recommender Systems for Negative Emotion Handling in Therapy Chatbots18 Nov 2023 0 repositories listed
-
A Study on Altering the Latent Space of Pretrained Text to Speech Models for Improved Expressiveness17 Nov 2023 0 repositories listed
-
ChatAnything: Facetime Chat with LLM-Enhanced Personas12 Nov 2023 0 repositories listed
-
Synthetic Speaking Children -- Why We Need Them and How to Make Them8 Nov 2023 0 repositories listed
-
Character-Level Bangla Text-to-IPA Transcription Using Transformer Architecture with Sequence Alignment7 Nov 2023 0 repositories listed
-
Transduce and Speak: Neural Transducer for Text-to-Speech with Semantic Token Prediction6 Nov 2023 0 repositories listed
-
E3 TTS: Easy End-to-End Diffusion-based Text to Speech2 Nov 2023 0 repositories listed
-
Expressive TTS Driven by Natural Language Prompts Using Few Human Annotations2 Nov 2023 0 repositories listed
-
Style Description based Text-to-Speech with Conditional Prosodic Layer Normalization based Diffusion GAN27 Oct 2023 0 repositories listed
-
Generative Pre-training for Speech with Flow Matching25 Oct 2023 0 repositories listed
-
DPP-TTS: Diversifying prosodic features of speech via determinantal point processes23 Oct 2023 0 repositories listed
-
An overview of text-to-speech systems and media applications22 Oct 2023 0 repositories listed
-
On the Relevance of Phoneme Duration Variability of Synthesized Training Data for Automatic Speech Recognition12 Oct 2023 0 repositories listed
-
Comparative Analysis of Transfer Learning in Deep Learning Text-to-Speech Models on a Few-Shot, Low-Resource, Customized Dataset8 Oct 2023 0 repositories listed
-
8 Oct 2023 0 repositories listed
-
8 Oct 2023 0 repositories listed
-
Latent Filling: Latent Space Data Augmentation for Zero-shot Speech Synthesis5 Oct 2023 0 repositories listed
-
The VoiceMOS Challenge 2023: Zero-shot Subjective Speech Quality Prediction for Multiple Domains4 Oct 2023 0 repositories listed
-
Towards human-like spoken dialogue generation between AI agents from written dialogue2 Oct 2023 0 repositories listed
-
Low-Resource Self-Supervised Learning with SSL-Enhanced TTS29 Sep 2023 0 repositories listed
-
Synthetic Speech Detection Based on Temporal Consistency and Distribution of Speaker Features29 Sep 2023 0 repositories listed
-
High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion Models27 Sep 2023 0 repositories listed
-
Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping25 Sep 2023 0 repositories listed
-
VoiceLDM: Text-to-Speech with Environmental Context24 Sep 2023 0 repositories listed
-
DurIAN-E: Duration Informed Attention Network For Expressive Text-to-Speech Synthesis22 Sep 2023 0 repositories listed
-
The Impact of Silence on Speech Anti-Spoofing21 Sep 2023 0 repositories listed
-
Speak While You Think: Streaming Speech Synthesis During Text Generation20 Sep 2023 0 repositories listed
-
Exploring Speech Enhancement for Low-resource Speech Synthesis19 Sep 2023 0 repositories listed
-
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition19 Sep 2023 0 repositories listed
-
Augmenting text for spoken language understanding with Large Language Models17 Sep 2023 0 repositories listed
-
Cross-lingual Knowledge Distillation via Flow-based Voice Conversion for Robust Polyglot Text-To-Speech15 Sep 2023 0 repositories listed
-
PromptTTS++: Controlling Speaker Identity in Prompt-Based Text-to-Speech Using Natural Language Descriptions15 Sep 2023 0 repositories listed
-
Direct Text to Speech Translation System using Acoustic Units14 Sep 2023 0 repositories listed
-
Cross-Utterance Conditioned VAE for Speech Generation8 Sep 2023 0 repositories listed
-
Large-Scale Automatic Audiobook Creation7 Sep 2023 0 repositories listed
-
GRASS: Unified Generation Model for Speech-to-Semantic Tasks6 Sep 2023 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.