Papers › Learning Speaker Embedding from Text-to-Speech

Learning Speaker Embedding from Text-to-Speech

21 Oct 2020arXiv:2010.11221archive 2025-07-28

Jaejin Cho, Piotr Zelasko, Jesus Villalba, Shinji Watanabe, Najim Dehak

Zero-shot multi-speaker Text-to-Speech (TTS) generates target speaker voices given an input text and the corresponding speaker embedding. In this work, we investigate the effectiveness of the TTS reconstruction objective to improve representation learning for speaker verification. We jointly trained end-to-end Tacotron 2 TTS and speaker embedding networks in a self-supervised fashion. We hypothesize that the embeddings will contain minimal phonetic information since the TTS decoder will obtain that information from the textual input. TTS reconstruction can also be combined with speaker classification to enhance these embeddings further. Once trained, the speaker encoder computes representations for the speaker verification task, while the rest of the TTS blocks are discarded. We investigated training TTS from either manual or ASR-generated transcripts. The latter allows us to train embeddings on datasets without manual transcripts. We compared ASR transcripts and Kaldi phone alignments as TTS inputs, showing that the latter performed better due to their finer resolution. Unsupervised TTS embeddings improved EER by 2.06\% absolute with regard to i-vectors for the LibriTTS dataset. TTS with speaker classification loss improved EER by 0.28\% and 0.73\% absolutely from a model using only speaker classification loss in LibriTTS and Voxceleb1 respectively.

PaperPDFCode

Code

JaejinCho/espnet_spkidtts officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClassificationDecoderGeneral ClassificationRepresentation LearningSpeaker VerificationText to Speechtext-to-speech

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

Batch NormalizationBiGRUBiLSTMCBHGConvolutionDense ConnectionsDilated Causal ConvolutionDropoutGRUGriffin-Lim AlgorithmHighway LayerHighway NetworkLSTMLinear LayerLocation Sensitive AttentionMax PoolingMixture of Logistic DistributionsReLUResidual ConnectionResidual GRUSigmoid ActivationTacotronTacotron 2Tanh ActivationWaveNetZoneout

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections