Papers › One Model, Many Languages: Meta-learning for Multilingual Text-to-Speech

One Model, Many Languages: Meta-learning for Multilingual Text-to-Speech

3 Aug 2020arXiv:2008.00768archive 2025-07-28

Tomáš Nekvinda, Ondřej Dušek

We introduce an approach to multilingual speech synthesis which uses the meta-learning concept of contextual parameter generation and produces natural-sounding multilingual speech using more languages and less training data than previous approaches. Our model is based on Tacotron 2 with a fully convolutional input text encoder whose weights are predicted by a separate parameter generator network. To boost voice cloning, the model uses an adversarial speaker classifier with a gradient reversal layer that removes speaker-specific information from the encoder. We arranged two experiments to compare our model with baselines using various levels of cross-lingual parameter sharing, in order to evaluate: (1) stability and performance when training on low amounts of data, (2) pronunciation accuracy and voice quality of code-switching synthesis. For training, we used the CSS10 dataset and our new small dataset based on Common Voice recordings in five languages. Our model is shown to effectively share information across languages and according to a subjective evaluation test, it produces more natural and accurate code-switching speech than the baselines.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Tomiinek/Multilingual_Text_to_Speech officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Meta-LearningSpeech SynthesisText to SpeechVoice Cloningtext-to-speech

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

Batch NormalizationBiGRUBiLSTMCBHGConvolutionDense ConnectionsDilated Causal ConvolutionDropoutGRUGriffin-Lim AlgorithmHighway LayerHighway NetworkLSTMLinear LayerLocation Sensitive AttentionMax PoolingMixture of Logistic DistributionsReLUResidual ConnectionResidual GRUSigmoid ActivationTacotronTacotron 2Tanh ActivationWaveNetZoneout

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections