Browse State-of-the-Art › Text to Speech › Papers, page 11
Text to Speech
Papers archive 2025-07-28
archive papers tagged: 1,419 · with a code link: 399 · where Syntology ran a sample: 108 (96 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (108 of 1,419 tagged: 96 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument)
Page 11 of 15: papers 1,001 to 1,100 of 1,419, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Building Open-source Speech Technology for Low-resource Minority Languages with SáMi as an Example – Tools, Methods and Experiments1 Jun 2022 0 repositories listed
-
Error Annotation in Post-Editing Machine Translation: Investigating the Impact of Text-to-Speech Technology1 Jun 2022 0 repositories listed
-
Exploring Transfer Learning for Urdu Speech Synthesis1 Jun 2022 0 repositories listed
-
Huqariq: A Multilingual Speech Corpus of Native Languages of Peru forSpeech Recognition1 Jun 2022 0 repositories listed
-
Investigating Inter- and Intra-speaker Voice Conversion using Audiobooks1 Jun 2022 0 repositories listed
-
ParlamentParla: A Speech Corpus of Catalan Parliamentary Sessions1 Jun 2022 0 repositories listed
-
Reading Assistance through LARA, the Learning And Reading Assistant1 Jun 2022 0 repositories listed
-
Text-to-Speech for Under-Resourced Languages: Phoneme Mapping and Source Language Selection in Transfer Learning1 Jun 2022 0 repositories listed
-
The Nós Project: Opening routes for the Galician language in the field of language technologies1 Jun 2022 0 repositories listed
-
Using the LARA Little Prince to compare human and TTS audio quality1 Jun 2022 0 repositories listed
-
Guided-TTS 2: A Diffusion Model for High-quality Adaptive Text-to-Speech with Untranscribed Data30 May 2022 0 repositories listed
-
Exploiting Transliterated Words for Finding Similarity in Inter-Language News Articles using Machine Learning29 May 2022 0 repositories listed
-
T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine Translation24 May 2022 0 repositories listed
-
ReCAB-VAE: Gumbel-Softmax Variational Inference Based on Analytic Divergence9 May 2022 0 repositories listed
-
Regotron: Regularizing the Tacotron2 architecture via monotonic alignment loss28 Apr 2022 0 repositories listed
-
Cross-Speaker Emotion Transfer for Low-Resource Text-to-Speech Using Non-Parallel Voice Conversion with Pitch-Shift Data Augmentation21 Apr 2022 0 repositories listed
-
Audio Deep Fake Detection System with Neural Stitching for ADD 202219 Apr 2022 0 repositories listed
-
Applying Feature Underspecified Lexicon Phonological Features in Multilingual Text-to-Speech14 Apr 2022 0 repositories listed
-
Study of Indian English Pronunciation Variabilities relative to Received Pronunciation13 Apr 2022 0 repositories listed
-
Enhancement of Pitch Controllability using Timbre-Preserving Pitch Augmentation in FastPitch12 Apr 2022 0 repositories listed
-
Fine-grained Noise Control for Multispeaker Speech Synthesis11 Apr 2022 0 repositories listed
-
The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance11 Apr 2022 0 repositories listed
-
Hierarchical and Multi-Scale Variational Autoencoder for Diverse and Natural Non-Autoregressive Text-to-Speech8 Apr 2022 0 repositories listed
-
Karaoker: Alignment-free singing voice synthesis with speech training data8 Apr 2022 0 repositories listed
-
Arabic Text-To-Speech (TTS) Data Preparation7 Apr 2022 0 repositories listed
-
Unsupervised Quantized Prosody Representation for Controllable Speech Synthesis7 Apr 2022 0 repositories listed
-
Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation6 Apr 2022 0 repositories listed
-
Representation Selective Self-distillation and wav2vec 2.0 Feature Exploration for Spoof-aware Speaker Verification6 Apr 2022 0 repositories listed
-
6 Apr 2022 0 repositories listed
-
Anti-Spoofing Using Transfer Learning with Variational Information Bottleneck4 Apr 2022 0 repositories listed
-
Deliberation Model for On-Device Spoken Language Understanding4 Apr 2022 0 repositories listed
-
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature2 Apr 2022 0 repositories listed
-
AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios1 Apr 2022 0 repositories listed
-
Text-To-Speech Data Augmentation for Low Resource Speech Recognition1 Apr 2022 0 repositories listed
-
Effectiveness of text to speech pseudo labels for forced alignment and cross lingual pretrained models for low resource speech recognition31 Mar 2022 0 repositories listed
-
Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech31 Mar 2022 0 repositories listed
-
31 Mar 2022 0 repositories listed
-
WavThruVec: Latent speech representation as intermediate features for neural speech synthesis31 Mar 2022 0 repositories listed
-
Does Audio Deepfake Detection Generalize?30 Mar 2022 0 repositories listed
-
Applying Syntax–Prosody Mapping Hypothesis and Prosodic Well-Formedness Constraints to Neural Sequence-to-Sequence Speech Synthesis29 Mar 2022 0 repositories listed
-
Transfer Learning Framework for Low-Resource Text-to-Speech using a Large-Scale Unlabeled Speech Corpus29 Mar 2022 0 repositories listed
-
STUDIES: Corpus of Japanese Empathetic Dialogue Speech Towards Friendly Voice Agent28 Mar 2022 0 repositories listed
-
Bunched LPCNet2: Efficient Neural Vocoders Covering Devices from Cloud to Edge27 Mar 2022 0 repositories listed
-
A Text-to-Speech Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results for Child Speech Synthesis22 Mar 2022 0 repositories listed
-
AutoTTS: End-to-End Text-to-Speech Synthesis through Differentiable Duration Modeling21 Mar 2022 0 repositories listed
-
Vocal effort modeling in neural TTS for improving the intelligibility of synthetic speech in noise20 Mar 2022 0 repositories listed
-
Improve few-shot voice cloning using multi-modal learning18 Mar 2022 0 repositories listed
-
Text-free non-parallel many-to-many voice conversion using normalising flows15 Mar 2022 0 repositories listed
-
Revisiting Over-Smoothness in Text to Speech26 Feb 2022 0 repositories listed
-
Human Detection of Political Speech Deepfakes across Transcripts, Audio, and Video25 Feb 2022 0 repositories listed
-
Improving Cross-lingual Speech Synthesis with Triplet Training Scheme22 Feb 2022 0 repositories listed
-
r-G2P: Evaluating and Enhancing Robustness of Grapheme to Phoneme Conversion by Controlled noise introducing and Contextual information incorporation21 Feb 2022 0 repositories listed
-
ProsoSpeech: Enhancing Prosody With Quantized Vector Pre-training in Text-to-Speech16 Feb 2022 0 repositories listed
-
Voice Filter: Few-shot text-to-speech speaker adaptation using voice conversion as a post-processing module16 Feb 2022 0 repositories listed
-
NewsPod: Automatic and Interactive News Podcasts15 Feb 2022 0 repositories listed
-
Unsupervised word-level prosody tagging for controllable speech synthesis15 Feb 2022 0 repositories listed
-
Distribution augmentation for low-resource expressive text-to-speech13 Feb 2022 0 repositories listed
-
Deep Performer: Score-to-Audio Music Performance Synthesis12 Feb 2022 0 repositories listed
-
Cross-speaker style transfer for text-to-speech using data augmentation10 Feb 2022 0 repositories listed
-
Building Synthetic Speaker Profiles in Text-to-Speech Systems7 Feb 2022 0 repositories listed
-
Multi-Stage Deep Transfer Learning for EmIoT-enabled Human-Computer Interaction3 Feb 2022 0 repositories listed
-
Transformer-based Models of Text Normalization for Speech Applications1 Feb 2022 0 repositories listed
-
Synthesizing Dysarthric Speech Using Multi-talker TTS for Dysarthric Speech Recognition27 Jan 2022 0 repositories listed
-
The MSXF TTS System for ICASSP 2022 ADD Challenge27 Jan 2022 0 repositories listed
-
Zero-Shot Long-Form Voice Cloning with Dynamic Convolution Attention25 Jan 2022 0 repositories listed
-
Polyphone disambiguation and accent prediction using pre-trained language models in Japanese TTS front-end24 Jan 2022 0 repositories listed
-
Cross-Lingual Text-to-Speech Using Multi-Task Learning and Speaker Classifier Joint Training20 Jan 2022 0 repositories listed
-
Empathic Machines: Using Intermediate Features as Levers to Emulate Emotions in Text-To-Speech Systems16 Jan 2022 0 repositories listed
-
KazakhTTS2: Extending the Open-Source Kazakh TTS Corpus With More Data, Speakers, and Topics15 Jan 2022 0 repositories listed
-
SoK: A Study of the Security on Voice Processing Systems24 Dec 2021 0 repositories listed
-
Multi-speaker Multi-style Text-to-speech Synthesis With Single-speaker Single-style Training Data Scenarios23 Dec 2021 0 repositories listed
-
Multi-speaker Emotional Text-to-speech Synthesizer7 Dec 2021 0 repositories listed
-
Speech-T: Transducer for Text to Speech and Beyond1 Dec 2021 0 repositories listed
-
Generating Rich Product Descriptions for Conversational E-commerce Systems30 Nov 2021 0 repositories listed
-
Guided-TTS: A Diffusion Model for Text-to-Speech via Classifier Guidance23 Nov 2021 0 repositories listed
-
Improved Prosodic Clustering for Multispeaker and Speaker-independent Phoneme-level Prosody Control19 Nov 2021 0 repositories listed
-
Prosodic Clustering for Phoneme-level Prosody Control in End-to-End Speech Synthesis19 Nov 2021 0 repositories listed
-
Semi-supervised transfer learning for language expansion of end-to-end speech recognition models to low-resource languages19 Nov 2021 0 repositories listed
-
High Quality Streaming Speech Synthesis with Low, Sentence-Length-Independent Latency17 Nov 2021 0 repositories listed
-
Cross-Utterance Conditioned VAE for Non-Autoregressive Text-to-Speech16 Nov 2021 0 repositories listed
-
Speech Synthesis for Low Resource Languages using Transliteration Enabled Transfer Learning16 Nov 2021 0 repositories listed
-
Meta-Voice: Fast few-shot style transfer for expressive voice cloning using meta learning14 Nov 2021 0 repositories listed
-
Emotional Prosody Control for Speech Generation7 Nov 2021 0 repositories listed
-
Speaker Generation7 Nov 2021 0 repositories listed
-
Controlling Prosody in End-to-End TTS: A Case Study on Contrastive Focus Generation1 Nov 2021 0 repositories listed
-
ViDA-MAN: Visual Dialog with Digital Humans26 Oct 2021 0 repositories listed
-
Discrete Acoustic Space for an Efficient Sampling in Neural Text-To-Speech24 Oct 2021 0 repositories listed
-
From Start to Finish: Latency Reduction Strategies for Incremental Speech Synthesis in Simultaneous Speech-to-Speech Translation15 Oct 2021 0 repositories listed
-
Neural Dubber: Dubbing for Videos According to Scripts15 Oct 2021 0 repositories listed
-
Exploring Timbre Disentanglement in Non-Autoregressive Cross-Lingual Text-to-Speech14 Oct 2021 0 repositories listed
-
FedSpeech: Federated Text-to-Speech with Continual Learning14 Oct 2021 0 repositories listed
-
Improve Cross-lingual Voice Cloning Using Low-quality Code-switched Data14 Oct 2021 0 repositories listed
-
Revisiting IPA-based Cross-lingual Text-to-speech14 Oct 2021 0 repositories listed
-
SingGAN: Generative Adversarial Network For High-Fidelity Singing Voice Generation14 Oct 2021 0 repositories listed
-
A Melody-Unsupervision Model for Singing Voice Synthesis13 Oct 2021 0 repositories listed
-
Adapting TTS models For New Speakers using Transfer Learning12 Oct 2021 0 repositories listed
-
A study on the efficacy of model pre-training in developing neural text-to-speech system8 Oct 2021 0 repositories listed
-
Environment Aware Text-to-Speech Synthesis8 Oct 2021 0 repositories listed
-
VisualTTS: TTS with Accurate Lip-Speech Synchronization for Automatic Voice Over7 Oct 2021 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.