Browse State-of-the-Art › Speech Synthesis › Papers, page 6
Speech Synthesis
Papers archive 2025-07-28
archive papers tagged: 1,249 · with a code link: 366 · where Syntology ran a sample: 101 (85 with a run with no instrument failure, 16 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (101 of 1,249 tagged: 85 with a run with no instrument failure, 16 where every run was a failure of Syntology's instrument)
Page 6 of 13: papers 501 to 600 of 1,249, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Which Prosodic Features Matter Most for Pragmatics?23 Aug 2024 0 repositories listed
-
AI-Based IVR20 Aug 2024 0 repositories listed
-
VNet: A GAN-based Multi-Tier Discriminator Network for Speech Synthesis Vocoders13 Aug 2024 0 repositories listed
-
Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation1 Aug 2024 0 repositories listed
-
Speech Bandwidth Expansion Via High Fidelity Generative Adversarial Networks26 Jul 2024 0 repositories listed
-
Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models26 Jul 2024 0 repositories listed
-
Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning21 Jul 2024 0 repositories listed
-
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis19 Jul 2024 0 repositories listed
-
Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models18 Jul 2024 0 repositories listed
-
Autoregressive Speech Synthesis without Vector Quantization11 Jul 2024 0 repositories listed
-
Toward accessible comics for blind and low vision readers11 Jul 2024 0 repositories listed
-
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation8 Jul 2024 0 repositories listed
-
FA-GAN: Artifacts-free and Phase-aware High-fidelity GAN-based Vocoder5 Jul 2024 0 repositories listed
-
We Need Variations in Speech Generation: Sub-center Modelling for Speaker Embeddings5 Jul 2024 0 repositories listed
-
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis4 Jul 2024 0 repositories listed
-
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization2 Jul 2024 0 repositories listed
-
A Comprehensive Survey on Diffusion Models and Their Applications1 Jul 2024 0 repositories listed
-
Lightweight Zero-shot Text-to-Speech with Mixture of Adapters1 Jul 2024 0 repositories listed
-
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis30 Jun 2024 0 repositories listed
-
High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model25 Jun 2024 0 repositories listed
-
Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment25 Jun 2024 0 repositories listed
-
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation25 Jun 2024 0 repositories listed
-
Towards Zero-Shot Text-To-Speech for Arabic Dialects24 Jun 2024 0 repositories listed
-
A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge22 Jun 2024 0 repositories listed
-
A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis18 Jun 2024 0 repositories listed
-
1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis17 Jun 2024 0 repositories listed
-
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis16 Jun 2024 0 repositories listed
-
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis13 Jun 2024 0 repositories listed
-
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models12 Jun 2024 0 repositories listed
-
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment12 Jun 2024 0 repositories listed
-
Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?11 Jun 2024 0 repositories listed
-
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems11 Jun 2024 0 repositories listed
-
JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis10 Jun 2024 0 repositories listed
-
Text-aware and Context-aware Expressive Audiobook Speech Synthesis9 Jun 2024 0 repositories listed
-
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis8 Jun 2024 0 repositories listed
-
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers8 Jun 2024 0 repositories listed
-
Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs7 Jun 2024 0 repositories listed
-
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model6 Jun 2024 0 repositories listed
-
Style Mixture of Experts for Expressive Text-To-Speech Synthesis5 Jun 2024 0 repositories listed
-
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis4 Jun 2024 0 repositories listed
-
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training3 Jun 2024 0 repositories listed
-
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback2 Jun 2024 0 repositories listed
-
Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning23 May 2024 0 repositories listed
-
DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models23 May 2024 0 repositories listed
-
Evaluating Text-to-Speech Synthesis from a Large Discrete Token-based Speech Language Model16 May 2024 0 repositories listed
-
Expressivity and Speech Synthesis30 Apr 2024 0 repositories listed
-
Retrieval-Augmented Audio Deepfake Detection22 Apr 2024 0 repositories listed
-
Parameter Efficient Fine Tuning: A Comprehensive Analysis Across Applications21 Apr 2024 0 repositories listed
-
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis4 Apr 2024 0 repositories listed
-
Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation3 Apr 2024 0 repositories listed
-
PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders3 Apr 2024 0 repositories listed
-
Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling1 Apr 2024 0 repositories listed
-
Training Generative Adversarial Network-Based Vocoder with Limited Data Using Augmentation-Conditional Discriminator25 Mar 2024 0 repositories listed
-
21 Mar 2024 0 repositories listed
-
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis19 Mar 2024 0 repositories listed
-
EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech13 Mar 2024 0 repositories listed
-
Towards Accurate Lip-to-Speech Synthesis in-the-Wild2 Mar 2024 0 repositories listed
-
VoxGenesis: Unsupervised Discovery of Latent Speaker Manifold for Speech Synthesis1 Mar 2024 0 repositories listed
-
Extending Multilingual Speech Synthesis to 100+ Languages without Transcribed Data29 Feb 2024 0 repositories listed
-
16 Feb 2024 0 repositories listed Syntology 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 3 harvested samples) · 3 pointer-only (licence)
-
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis11 Feb 2024 0 repositories listed
-
SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition31 Jan 2024 0 repositories listed
-
SpecDiff-GAN: A Spectrally-Shaped Noise Diffusion GAN for Speech and Music Synthesis30 Jan 2024 0 repositories listed
-
MunTTS: A Text-to-Speech System for Mundari28 Jan 2024 0 repositories listed
-
Empowering Communication: Speech Technology for Indian and Western Accents through AI-powered Speech Synthesis22 Jan 2024 0 repositories listed
-
Ultra-lightweight Neural Differential DSP Vocoder For High Quality Speech Synthesis19 Jan 2024 0 repositories listed
-
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis16 Jan 2024 0 repositories listed
-
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters10 Jan 2024 0 repositories listed
-
StreamVC: Real-Time Low-Latency Voice Conversion5 Jan 2024 0 repositories listed
-
Incremental FastPitch: Chunk-based High Quality Text to Speech3 Jan 2024 0 repositories listed
-
Boosting Large Language Model for Speech Synthesis: An Empirical Study30 Dec 2023 0 repositories listed
-
Normalization of Lithuanian Text Using Regular Expressions29 Dec 2023 0 repositories listed
-
Creating New Voices using Normalizing Flows22 Dec 2023 0 repositories listed
-
BrainTalker: Low-Resource Brain-to-Speech Synthesis with Transfer Learning using Wav2Vec 2.021 Dec 2023 0 repositories listed
-
Evaluating Speech-in-Speech Perception via a Humanoid Robot19 Dec 2023 0 repositories listed
-
StyleSpeech: Self-supervised Style Enhancing with VQ-VAE-based Pre-training for Expressive Audiobook Speech Synthesis19 Dec 2023 0 repositories listed
-
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis17 Dec 2023 0 repositories listed
-
CONCSS: Contrastive-based Context Comprehension for Dialogue-appropriate Prosody in Conversational Speech Synthesis16 Dec 2023 0 repositories listed
-
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks10 Dec 2023 0 repositories listed
-
An Experimental Study: Assessing the Combined Framework of WavLM and BEST-RQ for Text-to-Speech Synthesis8 Dec 2023 0 repositories listed
-
Schrodinger Bridges Beat Diffusion Models on Text-to-Speech Synthesis6 Dec 2023 0 repositories listed
-
Code-Mixed Text to Speech Synthesis under Low-Resource Constraints2 Dec 2023 0 repositories listed
-
Guided Flows for Generative Modeling and Decision Making22 Nov 2023 0 repositories listed
-
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis20 Nov 2023 0 repositories listed
-
LE-SSL-MOS: Self-Supervised Learning MOS Prediction with Listener Enhancement17 Nov 2023 0 repositories listed
-
On the Opportunities of Green Computing: A Survey1 Nov 2023 0 repositories listed
-
Controllable Generation of Artificial Speaker Embeddings through Discovery of Principal Directions26 Oct 2023 0 repositories listed
-
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning26 Oct 2023 0 repositories listed
-
Generative Pre-training for Speech with Flow Matching25 Oct 2023 0 repositories listed
-
Energy-Based Models For Speech Synthesis19 Oct 2023 0 repositories listed
-
SelfVC: Voice Conversion With Iterative Refinement using Self Transformations14 Oct 2023 0 repositories listed
-
Speaking rate attention-based duration prediction for speed control TTS13 Oct 2023 0 repositories listed
-
Privacy-oriented manipulation of speaker representations10 Oct 2023 0 repositories listed
-
8 Oct 2023 0 repositories listed
-
Latent Filling: Latent Space Data Augmentation for Zero-shot Speech Synthesis5 Oct 2023 0 repositories listed
-
The VoiceMOS Challenge 2023: Zero-shot Subjective Speech Quality Prediction for Multiple Domains4 Oct 2023 0 repositories listed
-
High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion Models27 Sep 2023 0 repositories listed
-
Collaborative Watermarking for Adversarial Speech Synthesis26 Sep 2023 0 repositories listed
-
Face-StyleSpeech: Enhancing Zero-shot Speech Synthesis from Face Images with Improved Face-to-Speech Mapping25 Sep 2023 0 repositories listed
-
DurIAN-E: Duration Informed Attention Network For Expressive Text-to-Speech Synthesis22 Sep 2023 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.