Browse State-of-the-Art › text-to-speech › Papers, page 10
text-to-speech
Papers archive 2025-07-28
archive papers tagged: 1,413 · with a code link: 395 · where Syntology ran a sample: 106 (95 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (106 of 1,413 tagged: 95 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument)
Page 10 of 15: papers 901 to 1,000 of 1,413, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Using External Off-Policy Speech-To-Text Mappings in Contextual End-To-End Automated Speech Recognition6 Jan 2023 0 repositories listed
-
ReVISE: Self-Supervised Speech Resynthesis With Visual Input for Universal and Generalized Speech Regeneration1 Jan 2023 0 repositories listed
-
HMM-based data augmentation for E2E systems for building conversational speech synthesis systems22 Dec 2022 0 repositories listed
-
21 Dec 2022 0 repositories listed
-
Improving the quality of neural TTS using long-form content and multi-speaker multi-style modeling20 Dec 2022 0 repositories listed
-
TTS-Guided Training for Accent Conversion Without Parallel Data20 Dec 2022 0 repositories listed
-
Investigation of Japanese PnG BERT language model in text-to-speech synthesis for pitch accent language16 Dec 2022 0 repositories listed
-
Speech Aware Dialog System Technology Challenge (DSTC11)16 Dec 2022 0 repositories listed
-
Text-to-speech synthesis based on latent variable conversion using diffusion probabilistic model and variational autoencoder16 Dec 2022 0 repositories listed
-
Probing Deep Speaker Embeddings for Speaker-related Tasks14 Dec 2022 0 repositories listed
-
Analysis and Utilization of Entrainment on Acoustic and Emotion Features in User-agent Dialogue7 Dec 2022 0 repositories listed
-
Low-Resource End-to-end Sanskrit TTS using Tacotron2, WaveGlow and Transfer Learning7 Dec 2022 0 repositories listed
-
SNAC: Speaker-normalized affine coupling layer in flow-based architecture for zero-shot multi-speaker text-to-speech30 Nov 2022 0 repositories listed
-
Controllable speech synthesis by learning discrete phoneme-level prosodic representations29 Nov 2022 0 repositories listed
-
Evaluating and reducing the distance between synthetic and real speech distributions29 Nov 2022 0 repositories listed
-
Contextual Expressive Text-to-Speech26 Nov 2022 0 repositories listed
-
Efficient Incremental Text-to-Speech on GPUs25 Nov 2022 0 repositories listed
-
23 Nov 2022 0 repositories listed
-
Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models17 Nov 2022 0 repositories listed
-
Back-Translation-Style Data Augmentation for Mandarin Chinese Polyphone Disambiguation17 Nov 2022 0 repositories listed
-
EmoDiff: Intensity Controllable Emotional Text-to-Speech with Soft-Label Guidance17 Nov 2022 0 repositories listed
-
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech14 Nov 2022 0 repositories listed
-
Semi-supervised learning for continuous emotional intensity controllable speech synthesis with disentangled representations11 Nov 2022 0 repositories listed
-
An Empirical Study on L2 Accents of Cross-lingual Text-to-Speech Systems via Vowel Space6 Nov 2022 0 repositories listed
-
Parallel Attention Forcing for Machine Translation6 Nov 2022 0 repositories listed
-
Stutter-TTS: Controlled Synthesis and Improved Recognition of Stuttered Speech4 Nov 2022 0 repositories listed
-
Generating Multilingual Gender-Ambiguous Text-to-Speech Voices1 Nov 2022 0 repositories listed
-
Investigating Content-Aware Neural Text-To-Speech MOS Prediction Using Prosodic and Linguistic Features1 Nov 2022 0 repositories listed
-
Technology Pipeline for Large Scale Cross-Lingual Dubbing of Lecture Videos into Multiple Indian Languages1 Nov 2022 0 repositories listed
-
Combining Automatic Speaker Verification and Prosody Analysis for Synthetic Speech Detection31 Oct 2022 0 repositories listed
-
Cross-lingual Text-To-Speech with Flow-based Voice Conversion for Improved Pronunciation31 Oct 2022 0 repositories listed
-
Structured State Space Decoder for Speech Recognition and Synthesis31 Oct 2022 0 repositories listed
-
Period VITS: Variational Inference with Explicit Pitch Modeling for End-to-end Emotional Speech Synthesis28 Oct 2022 0 repositories listed
-
Residual Adapters for Few-Shot Text-to-Speech Speaker Adaptation28 Oct 2022 0 repositories listed
-
Towards zero-shot Text-based voice editing using acoustic context conditioning, utterance embeddings, and reference encoders28 Oct 2022 0 repositories listed
-
Explicit Intensity Control for Accented Text-to-speech27 Oct 2022 0 repositories listed
-
Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-To-Speech27 Oct 2022 0 repositories listed
-
Improving Speech-to-Speech Translation Through Unlabeled Text26 Oct 2022 0 repositories listed
-
Adapitch: Adaption Multi-Speaker Text-to-Speech Conditioned on Pitch Disentangling with Untranscribed Data25 Oct 2022 0 repositories listed
-
Semi-Supervised Learning Based on Reference Model for Low-resource TTS25 Oct 2022 0 repositories listed
-
Efficiently Trained Low-Resource Mongolian Text-to-Speech System Based On FullConv-TTS24 Oct 2022 0 repositories listed
-
Adaptive re-calibration of channel-wise features for Adversarial Audio Classification21 Oct 2022 0 repositories listed
-
LeVoice ASR Systems for the ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge14 Oct 2022 0 repositories listed
-
Pre-Avatar: An Automatic Presentation Generation Framework Leveraging Talking Avatar13 Oct 2022 0 repositories listed
-
Adversarial Speaker-Consistency Learning Using Untranscribed Speech Data for Zero-Shot Multi-Speaker Text-to-Speech12 Oct 2022 0 repositories listed
-
SQuId: Measuring Speech Naturalness in Many Languages12 Oct 2022 0 repositories listed
-
An Overview of Affective Speech Synthesis and Conversion in the Deep Learning Era6 Oct 2022 0 repositories listed
-
Unsupervised Multi-scale Expressive Speaking Style Modeling with Hierarchical Context Information for Audiobook Speech Synthesis1 Oct 2022 0 repositories listed
-
Multi-Task Adversarial Training Algorithm for Multi-Speaker Neural Text-to-Speech26 Sep 2022 0 repositories listed
-
Controllable Accented Text-to-Speech Synthesis22 Sep 2022 0 repositories listed
-
EPIC TTS Models: Empirical Pruning Investigations Characterizing Text-To-Speech Models22 Sep 2022 0 repositories listed
-
Using Rater and System Metadata to Explain Variance in the VoiceMOS Challenge 2022 Dataset14 Sep 2022 0 repositories listed
-
SANIP: Shopping Assistant and Navigation for the visually impaired8 Sep 2022 0 repositories listed
-
Non-Standard Vietnamese Word Detection and Normalization for Text-to-Speech7 Sep 2022 0 repositories listed
-
Improving Contextual Recognition of Rare Words with an Alternate Spelling Prediction Model2 Sep 2022 0 repositories listed
-
Towards MOOCs for Lipreading: Using Synthetic Talking Heads to Train Humans in Lipreading at Scale21 Aug 2022 0 repositories listed
-
Speech Synthesis with Mixed Emotions11 Aug 2022 0 repositories listed
-
A Study of Modeling Rising Intonation in Cantonese Neural Speech Synthesis3 Aug 2022 0 repositories listed
-
Low-data? No problem: low-resource, language-agnostic conversational text-to-speech via F0-conditioned data augmentation29 Jul 2022 0 repositories listed
-
Transplantation of Conversational Speaking Style with Interjections in Sequence-to-Sequence Speech Synthesis25 Jul 2022 0 repositories listed
-
A Cyclical Approach to Synthetic and Natural Speech Mismatch Refinement of Neural Post-filter for Low-cost Text-to-speech System13 Jul 2022 0 repositories listed
-
SATTS: Speaker Attractor Text to Speech, Learning to Speak by Learning to Separate13 Jul 2022 0 repositories listed
-
Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS13 Jul 2022 0 repositories listed
-
End-to-end speech recognition modeling from de-identified data12 Jul 2022 0 repositories listed
-
Huqariq: A Multilingual Speech Corpus of Native Languages of Peru for Speech Recognition12 Jul 2022 0 repositories listed
-
LIP: Lightweight Intelligent Preprocessor for meaningful text-to-speech11 Jul 2022 0 repositories listed
-
BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model4 Jul 2022 0 repositories listed
-
Mix and Match: An Empirical Study on Training Corpus Composition for Polyglot Text-To-Speech (TTS)4 Jul 2022 0 repositories listed
-
Unify and Conquer: How Phonetic Feature Representation Affects Polyglot Text-To-Speech (TTS)4 Jul 2022 0 repositories listed
-
Computer-assisted Pronunciation Training -- Speech synthesis is almost all you need2 Jul 2022 0 repositories listed
-
A Polyphone BERT for Polyphone Disambiguation in Mandarin Chinese1 Jul 2022 0 repositories listed
-
Automatic Evaluation of Speaker Similarity1 Jul 2022 0 repositories listed
-
Empathic Machines: Using Intermediate Features as Levers to Emulate Emotions in Text-To-Speech Systems1 Jul 2022 0 repositories listed
-
Fast Bilingual Grapheme-To-Phoneme Conversion1 Jul 2022 0 repositories listed
-
R-MelNet: Reduced Mel-Spectral Modeling for Neural TTS30 Jun 2022 0 repositories listed
-
TTS-by-TTS 2: Data-selective augmentation for neural speech synthesis using ranking support vector machine with variational autoencoder30 Jun 2022 0 repositories listed
-
Improving Deliberation by Text-Only and Semi-Supervised Training29 Jun 2022 0 repositories listed
-
Simple and Effective Multi-sentence TTS with Expressive and Coherent Prosody29 Jun 2022 0 repositories listed
-
Comparison of Speech Representations for the MOS Prediction System28 Jun 2022 0 repositories listed
-
Expressive, Variable, and Controllable Duration Modelling in TTS28 Jun 2022 0 repositories listed
-
Few-Shot Cross-Lingual TTS Using Transferable Phoneme Embedding27 Jun 2022 0 repositories listed
-
Synthesizing Personalized Non-speech Vocalization from Discrete Speech Representations25 Jun 2022 0 repositories listed
-
End-to-End Text-to-Speech Based on Latent Representation of Speaking Styles Using Spontaneous Dialogue24 Jun 2022 0 repositories listed
-
SANE-TTS: Stable And Natural End-to-End Multilingual Text-to-Speech24 Jun 2022 0 repositories listed
-
A Simple Baseline for Domain Adaptation in End to End ASR Systems Using Synthetic Data22 Jun 2022 0 repositories listed
-
Human-in-the-loop Speaker Adaptation for DNN-based Multi-speaker TTS21 Jun 2022 0 repositories listed
-
Towards Optimizing OCR for Accessibility21 Jun 2022 0 repositories listed
-
NatiQ: An End-to-end Text-to-Speech System for Arabic15 Jun 2022 0 repositories listed
-
A Novel Chinese Dialect TTS Frontend with Non-Autoregressive Neural Machine Translation10 Jun 2022 0 repositories listed
-
Face-Dubbing++: Lip-Synchronous, Voice Preserving Translation of Videos9 Jun 2022 0 repositories listed
-
FlexLip: A Controllable Text-to-Lip System7 Jun 2022 0 repositories listed
-
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE6 Jun 2022 0 repositories listed
-
Audiobook Dialogues as Training Data for Conversational Style Synthetic Voices1 Jun 2022 0 repositories listed
-
BU-TTS: An Open-Source, Bilingual Welsh-English, Text-to-Speech Corpus1 Jun 2022 0 repositories listed
-
Building Open-source Speech Technology for Low-resource Minority Languages with SáMi as an Example – Tools, Methods and Experiments1 Jun 2022 0 repositories listed
-
Error Annotation in Post-Editing Machine Translation: Investigating the Impact of Text-to-Speech Technology1 Jun 2022 0 repositories listed
-
Exploring Transfer Learning for Urdu Speech Synthesis1 Jun 2022 0 repositories listed
-
Huqariq: A Multilingual Speech Corpus of Native Languages of Peru forSpeech Recognition1 Jun 2022 0 repositories listed
-
Investigating Inter- and Intra-speaker Voice Conversion using Audiobooks1 Jun 2022 0 repositories listed
-
ParlamentParla: A Speech Corpus of Catalan Parliamentary Sessions1 Jun 2022 0 repositories listed