Browse State-of-the-Art › Speech Synthesis › Papers, page 8
Speech Synthesis
Papers archive 2025-07-28
archive papers tagged: 1,249 · with a code link: 366 · where Syntology ran a sample: 101 (85 with a run with no instrument failure, 16 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (101 of 1,249 tagged: 85 with a run with no instrument failure, 16 where every run was a failure of Syntology's instrument)
Page 8 of 13: papers 701 to 800 of 1,249, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models17 Nov 2022 0 repositories listed
-
Audio Anti-spoofing Using a Simple Attention Module and Joint Optimization Based on Additive Angular Margin Loss and Meta-learning17 Nov 2022 0 repositories listed
-
The Potential of Neural Speech Synthesis-based Data Augmentation for Personalized Speech Enhancement14 Nov 2022 0 repositories listed
-
Semi-supervised learning for continuous emotional intensity controllable speech synthesis with disentangled representations11 Nov 2022 0 repositories listed
-
Deliberation Networks and How to Train Them6 Nov 2022 0 repositories listed
-
Predicting phoneme-level prosody latents using AR and flow-based Prior Networks for expressive speech synthesis2 Nov 2022 0 repositories listed
-
A Preliminary Study on Mandarin-Hakka neural machine translation using small-sized data1 Nov 2022 0 repositories listed
-
Development of Mandarin-English code-switching speech synthesis system1 Nov 2022 0 repositories listed
-
Learning utterance-level representations through token-level acoustic latents prediction for Expressive Speech Synthesis1 Nov 2022 0 repositories listed
-
Taiwanese-Accented Mandarin and English Multi-Speaker Talking-Face Synthesis System1 Nov 2022 0 repositories listed
-
Technology Pipeline for Large Scale Cross-Lingual Dubbing of Lecture Videos into Multiple Indian Languages1 Nov 2022 0 repositories listed
-
Towards Developing State-of-the-Art TTS Synthesisers for 13 Indian Languages with Signal Processing aided Alignments31 Oct 2022 0 repositories listed
-
Period VITS: Variational Inference with Explicit Pitch Modeling for End-to-end Emotional Speech Synthesis28 Oct 2022 0 repositories listed
-
Virtuoso: Massive Multilingual Speech-Text Joint Semi-Supervised Learning for Text-To-Speech27 Oct 2022 0 repositories listed
-
Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks26 Oct 2022 0 repositories listed
-
RedPen: Region- and Reason-Annotated Dataset of Unnatural Speech26 Oct 2022 0 repositories listed
-
Semi-Supervised Learning Based on Reference Model for Low-resource TTS25 Oct 2022 0 repositories listed
-
A Data-Driven Investigation of Noise-Adaptive Utterance Generation with Linguistic Modification19 Oct 2022 0 repositories listed
-
Simple and Effective Unsupervised Speech Translation18 Oct 2022 0 repositories listed
-
Transformer-Based Speech Synthesizer Attribution in an Open Set Scenario14 Oct 2022 0 repositories listed
-
An Overview of Affective Speech Synthesis and Conversion in the Deep Learning Era6 Oct 2022 0 repositories listed
-
Fully Unsupervised Training of Few-shot Keyword Spotting6 Oct 2022 0 repositories listed
-
Unsupervised Multi-scale Expressive Speaking Style Modeling with Hierarchical Context Information for Audiobook Speech Synthesis1 Oct 2022 0 repositories listed
-
Controllable Accented Text-to-Speech Synthesis22 Sep 2022 0 repositories listed
-
EPIC TTS Models: Empirical Pruning Investigations Characterizing Text-To-Speech Models22 Sep 2022 0 repositories listed
-
An Initial study on Birdsong Re-synthesis Using Neural Vocoders21 Sep 2022 0 repositories listed
-
AutoLV: Automatic Lecture Video Generator19 Sep 2022 0 repositories listed
-
Decoupled Pronunciation and Prosody Modeling in Meta-Learning-Based Multilingual Speech Synthesis14 Sep 2022 0 repositories listed
-
Automated detection of pronunciation errors in non-native English speech employing deep learning13 Sep 2022 0 repositories listed
-
Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild1 Sep 2022 0 repositories listed
-
Audio Deepfake Attribution: An Initial Dataset and Investigation21 Aug 2022 0 repositories listed
-
Speech Synthesis with Mixed Emotions11 Aug 2022 0 repositories listed
-
A Study of Modeling Rising Intonation in Cantonese Neural Speech Synthesis3 Aug 2022 0 repositories listed
-
Transplantation of Conversational Speaking Style with Interjections in Sequence-to-Sequence Speech Synthesis25 Jul 2022 0 repositories listed
-
Controllable Data Generation by Deep Learning: A Review19 Jul 2022 0 repositories listed
-
PoeticTTS -- Controllable Poetry Reading for Literary Studies11 Jul 2022 0 repositories listed
-
End-to-End Binaural Speech Synthesis8 Jul 2022 0 repositories listed
-
BERT, can HE predict contrastive focus? Predicting and controlling prominence in neural TTS using a language model4 Jul 2022 0 repositories listed
-
Mix and Match: An Empirical Study on Training Corpus Composition for Polyglot Text-To-Speech (TTS)4 Jul 2022 0 repositories listed
-
Computer-assisted Pronunciation Training -- Speech synthesis is almost all you need2 Jul 2022 0 repositories listed
-
R-MelNet: Reduced Mel-Spectral Modeling for Neural TTS30 Jun 2022 0 repositories listed
-
TTS-by-TTS 2: Data-selective augmentation for neural speech synthesis using ranking support vector machine with variational autoencoder30 Jun 2022 0 repositories listed
-
iEmoTTS: Toward Robust Cross-Speaker Emotion Transfer and Control for Speech Synthesis based on Disentanglement between Prosody and Timbre29 Jun 2022 0 repositories listed
-
Expressive, Variable, and Controllable Duration Modelling in TTS28 Jun 2022 0 repositories listed
-
Self-supervised Context-aware Style Representation for Expressive Speech Synthesis25 Jun 2022 0 repositories listed
-
WOLONet: Wave Outlooker for Efficient and High Fidelity Speech Synthesis20 Jun 2022 0 repositories listed
-
Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue History16 Jun 2022 0 repositories listed
-
VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection15 Jun 2022 0 repositories listed
-
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE6 Jun 2022 0 repositories listed
-
Pronunciation Dictionary-Free Multilingual Speech Synthesis by Combining Unsupervised and Supervised Phonetic Representations2 Jun 2022 0 repositories listed
-
AiRO - an Interactive Learning Tool for Children at Risk of Dyslexia1 Jun 2022 0 repositories listed
-
BU-TTS: An Open-Source, Bilingual Welsh-English, Text-to-Speech Corpus1 Jun 2022 0 repositories listed
-
Building Open-source Speech Technology for Low-resource Minority Languages with SáMi as an Example – Tools, Methods and Experiments1 Jun 2022 0 repositories listed
-
Exploring Transfer Learning for Urdu Speech Synthesis1 Jun 2022 0 repositories listed
-
Investigating Inter- and Intra-speaker Voice Conversion using Audiobooks1 Jun 2022 0 repositories listed
-
SyntAct: A Synthesized Database of Basic Emotions1 Jun 2022 0 repositories listed
-
Macedonian Speech Synthesis for Assistive Technology Applications18 May 2022 0 repositories listed
-
ReCAB-VAE: Gumbel-Softmax Variational Inference Based on Analytic Divergence9 May 2022 0 repositories listed
-
Attentive activation function for improving end-to-end spoofing countermeasure systems3 May 2022 0 repositories listed
-
A Post Auto-regressive GAN Vocoder Focused on Spectrum Fracture12 Apr 2022 0 repositories listed
-
Fine-grained Noise Control for Multispeaker Speech Synthesis11 Apr 2022 0 repositories listed
-
The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance11 Apr 2022 0 repositories listed
-
DDOS: A MOS Prediction Framework utilizing Domain Adaptive Pre-training and Distribution of Opinion Scores7 Apr 2022 0 repositories listed
-
MAESTRO: Matched Speech Text Representations through Modality Matching7 Apr 2022 0 repositories listed
-
Self-supervised learning for robust voice cloning7 Apr 2022 0 repositories listed
-
Unsupervised Quantized Prosody Representation for Controllable Speech Synthesis7 Apr 2022 0 repositories listed
-
Simple and Effective Unsupervised Speech Synthesis6 Apr 2022 0 repositories listed
-
6 Apr 2022 0 repositories listed
-
A Comparison of Deep Learning MOS Predictors for Speech Synthesis Quality5 Apr 2022 0 repositories listed
-
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature2 Apr 2022 0 repositories listed
-
AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios1 Apr 2022 0 repositories listed
-
Residual-guided Personalized Speech Synthesis based on Face Image1 Apr 2022 0 repositories listed
-
WavThruVec: Latent speech representation as intermediate features for neural speech synthesis31 Mar 2022 0 repositories listed
-
Applying Syntax–Prosody Mapping Hypothesis and Prosodic Well-Formedness Constraints to Neural Sequence-to-Sequence Speech Synthesis29 Mar 2022 0 repositories listed
-
Analysis of Voice Conversion and Code-Switching Synthesis Using VQ-VAE28 Mar 2022 0 repositories listed
-
Towards Expressive Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis23 Mar 2022 0 repositories listed
-
A Text-to-Speech Pipeline, Evaluation Methodology, and Initial Fine-Tuning Results for Child Speech Synthesis22 Mar 2022 0 repositories listed
-
Modeling speech recognition and synthesis simultaneously: Encoding and decoding lexical and sublexical semantic information into speech with no direct access to speech data22 Mar 2022 0 repositories listed
-
AutoTTS: End-to-End Text-to-Speech Synthesis through Differentiable Duration Modeling21 Mar 2022 0 repositories listed
-
AdaVocoder: Adaptive Vocoder for Custom Voice18 Mar 2022 0 repositories listed
-
Robotic Speech Synthesis: Perspectives on Interactions, Scenarios, and Ethics17 Mar 2022 0 repositories listed
-
Whither the Priors for (Vocal) Interactivity?16 Mar 2022 0 repositories listed
-
Text-free non-parallel many-to-many voice conversion using normalising flows15 Mar 2022 0 repositories listed
-
Speaker Adaption with Intuitive Prosodic Features for Statistical Parametric Speech Synthesis2 Mar 2022 0 repositories listed
-
Improving Cross-lingual Speech Synthesis with Triplet Training Scheme22 Feb 2022 0 repositories listed
-
VCVTS: Multi-speaker Video-to-Speech synthesis via cross-modal knowledge transfer from voice conversion18 Feb 2022 0 repositories listed
-
Voice Filter: Few-shot text-to-speech speaker adaptation using voice conversion as a post-processing module16 Feb 2022 0 repositories listed
-
Unsupervised word-level prosody tagging for controllable speech synthesis15 Feb 2022 0 repositories listed
-
Partially Fake Audio Detection by Self-attention-based Fake Span Discovery14 Feb 2022 0 repositories listed
-
Deep Performer: Score-to-Audio Music Performance Synthesis12 Feb 2022 0 repositories listed
-
Transformer-based Models of Text Normalization for Speech Applications1 Feb 2022 0 repositories listed
-
Zero-Shot Long-Form Voice Cloning with Dynamic Convolution Attention25 Jan 2022 0 repositories listed
-
Cross-Lingual Text-to-Speech Using Multi-Task Learning and Speaker Classifier Joint Training20 Jan 2022 0 repositories listed
-
Deep Speech Synthesis from Articulatory Features16 Jan 2022 0 repositories listed
-
Quasi-Taylor Samplers for Diffusion Generative Models based on Ideal Derivatives26 Dec 2021 0 repositories listed
-
Multi-speaker Multi-style Text-to-speech Synthesis With Single-speaker Single-style Training Data Scenarios23 Dec 2021 0 repositories listed
-
整合語者嵌入向量與後置濾波器於提升個人化合成語音之語者相似度 (Incorporating Speaker Embedding and Post-Filter Network for Improving Speaker Similarity of Personalized Speech Synthesis System)1 Dec 2021 0 repositories listed
-
Guided-TTS: A Diffusion Model for Text-to-Speech via Classifier Guidance23 Nov 2021 0 repositories listed
-
Prosodic Clustering for Phoneme-level Prosody Control in End-to-End Speech Synthesis19 Nov 2021 0 repositories listed
-
Word-Level Style Control for Expressive, Non-attentive Speech Synthesis19 Nov 2021 0 repositories listed