{"url":"/task/voice-conversion","name":"Voice Conversion","slug":"voice-conversion","description_markdown":"I remember all the summer days\r\nDrinking wine in the sunshine\r\nI hope it never leaves\r\nAnd I remember all the summer nights\r\nStaring at you in the moonlight\r\nI hope you never leave 'cause baby\r\nYou're so good to me\r\nYou have all that all that I ever need\r\nIt's easy to love you\r\nSo easy to love you\r\nOoh you know it's true\r\nThe best part of being with you\r\nTo know you're with me\r\nIt's not so hard to say\r\nIt's easy to love you\r\nI remember all those winter days frozen\r\nIn the cold tryin' to get you home\r\nShould I be moving in, we can be together then\r\nRemember spending all those winter nights\r\nStayin' inside by the warm fire\r\nYeah you gotta know that I can never let you go\r\nYou and I have the rest of our lives to say\r\nIt's easy to love you\r\nSo easy to love you\r\nOoh you know it's true\r\nThe best part of being with you\r\nTo know you're with me\r\nIt's not so hard to say\r\nIt's easy to love you\r\nCan anybody else see it?\r\nMm, can anybody else see what I do?\r\nCan anybody else feel it?\r\nOh, can anybody else feel the way I do?\r\nBut now I'm with you\r\nHard to forget all the moments when\r\nWe'd be sitting there hoping it would never end\r\n'Cause this is meant to be\r\nSo baby, will you marry me?\r\nIt's easy to love you\r\nSo easy to love you\r\nOoh, you know it's true\r\nThe best part of being with you\r\nTo know you are with me\r\nIt's not so hard to say\r\nIt's easy to love you\r\nYou and me will be together\r\nI know our love will last forever\r\nYou and me will be together\r\nI know our love will last forever\r\nYou know it's true\r\nThe best part of being with you\r\nYou're easy to love\r\n\r\n<span class=\"description-source\">Source: [Joint training framework for text-to-speech and voice conversion using multi-source Tacotron and WaveNet ](https://arxiv.org/abs/1903.12389)</span>","categories":[{"name":"Audio","url":"/area/audio"},{"name":"Speech","url":"/area/speech"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":520,"papers_with_code":175,"benchmarks":3,"benchmark_tables_in_archive":3,"benchmark_tables_shown":3,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":8,"subtasks":0,"parent_tasks":2},"benchmarks":[{"leaderboard":"/sota/voice-conversion-on-zerospeech-2019-english","slug":"voice-conversion-on-zerospeech-2019-english","dataset":"ZeroSpeech 2019 English","dataset_url":null,"rows_in_archive":2,"metrics":["Speaker Similarity"],"first_row_in_archive_order":{"model":"VQ-CPC","paper_title":"Vector-quantized neural networks for acoustic unit discovery in the ZeroSpeech 2020 challenge","paper_url":"/paper/vector-quantized-neural-networks-for-acoustic","paper_date":"2020-05-19","arxiv_id":"2005.09409","code_links":[{"title":"bshall/ZeroSpeech","url":"https://github.com/bshall/ZeroSpeech"},{"title":"bshall/VectorQuantizedCPC","url":"https://github.com/bshall/VectorQuantizedCPC"}],"syntology":null}},{"leaderboard":"/sota/voice-conversion-on-librispeech-test-clean","slug":"voice-conversion-on-librispeech-test-clean","dataset":"LibriSpeech test-clean","dataset_url":"/dataset/librispeech","rows_in_archive":1,"metrics":["Character Error Rate (CER)","Equal Error Rate","Word Error Rate (WER)"],"first_row_in_archive_order":{"model":"kNN-VC (prematched HiFiGAN)","paper_title":"Voice Conversion With Just Nearest Neighbors","paper_url":"/paper/voice-conversion-with-just-nearest-neighbors","paper_date":"2023-05-30","arxiv_id":"2305.18975","code_links":[{"title":"bshall/knn-vc","url":"https://github.com/bshall/knn-vc"}],"syntology":null}},{"leaderboard":"/sota/voice-conversion-on-vctk","slug":"voice-conversion-on-vctk","dataset":"VCTK","dataset_url":"/dataset/vctk","rows_in_archive":1,"metrics":["Total Length Error (TLE)","Word Length Error (WLE)","Phone Length Error (PLE)"],"first_row_in_archive_order":{"model":"DISSC","paper_title":"Speaking Style Conversion in the Waveform Domain Using Discrete Self-Supervised Units","paper_url":"/paper/speaking-style-conversion-with-discrete-self","paper_date":"2022-12-19","arxiv_id":"2212.09730","code_links":[{"title":"gallilmaimon/DISSC","url":"https://github.com/gallilmaimon/DISSC"}],"syntology":{"n":7,"n_ran":3,"n_unverified":4,"n_pointer_only":0}}}],"datasets":[{"url":"/dataset/librispeech","name":"LibriSpeech","full_name":"","num_papers_in_archive":2361},{"url":"/dataset/vctk","name":"VCTK","full_name":"CSTR VCTK Corpus","num_papers_in_archive":476},{"url":"/dataset/esd","name":"ESD","full_name":"Emotional Speech Database","num_papers_in_archive":63},{"url":"/dataset/vivos","name":"VIVOS","full_name":"VIVOS Corpus","num_papers_in_archive":7},{"url":"/dataset/arvoice","name":"ArVoice","full_name":"ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis","num_papers_in_archive":1},{"url":"/dataset/gneutralspeech-female","name":"GneutralSpeech Female","full_name":"","num_papers_in_archive":1},{"url":"/dataset/gneutralspeech-male","name":"GneutralSpeech Male","full_name":"","num_papers_in_archive":1},{"url":"/dataset/vesus","name":"VESUS","full_name":"Varied Emotion in Syntactically Uniform Speech","num_papers_in_archive":1}],"subtasks":[],"parent_tasks":[{"url":"/task/1-image-2-2-stitchi","name":"1 Image, 2*2 Stitchi"},{"url":"/task/2d-classification","name":"2D Classification"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":175,"tagged_in_all":520,"items":[{"url":"/paper/stargan-vc-non-parallel-many-to-many-voice","title":"StarGAN-VC: Non-parallel many-to-many voice conversion with star generative adversarial networks","date":"2018-06-06","arxiv_id":"1806.02169","repositories_listed":14,"syntology":{"n":4,"n_ran":2,"n_unverified":2,"n_pointer_only":4}},{"url":"/paper/zero-shot-voice-style-transfer-with-only","title":"AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss","date":"2019-05-14","arxiv_id":"1905.05879","repositories_listed":11,"syntology":null},{"url":"/paper/one-shot-voice-conversion-by-separating","title":"One-shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization","date":"2019-04-10","arxiv_id":"1904.05742","repositories_listed":11,"syntology":{"n":9,"n_ran":0,"n_unverified":9,"n_pointer_only":0}},{"url":"/paper/parallel-data-free-voice-conversion-using","title":"Parallel-Data-Free Voice Conversion Using Cycle-Consistent Adversarial Networks","date":"2017-11-30","arxiv_id":"1711.11293","repositories_listed":9,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/utilizing-self-supervised-representations-for","title":"Utilizing Self-supervised Representations for MOS Prediction","date":"2021-04-07","arxiv_id":"2104.03017","repositories_listed":7,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/mosnet-deep-learning-based-objective","title":"MOSNet: Deep Learning based Objective Assessment for Voice Conversion","date":"2019-04-17","arxiv_id":"1904.08352","repositories_listed":7,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}},{"url":"/paper/speecht5-unified-modal-encoder-decoder-pre","title":"SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing","date":"2021-10-14","arxiv_id":"2110.07205","repositories_listed":6,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":3}},{"url":"/paper/defense-for-black-box-attacks-on-anti","title":"Defense for Black-box Attacks on Anti-spoofing Models by Self-Supervised Learning","date":"2020-06-05","arxiv_id":"2006.03214","repositories_listed":6,"syntology":null},{"url":"/paper/unsupervised-speech-decomposition-via-triple","title":"Unsupervised Speech Decomposition via Triple Information Bottleneck","date":"2020-04-23","arxiv_id":"2004.11284","repositories_listed":6,"syntology":{"n":12,"n_ran":0,"n_unverified":12,"n_pointer_only":0}},{"url":"/paper/cyclegan-vc2-improved-cyclegan-based-non","title":"CycleGAN-VC2: Improved CycleGAN-based Non-parallel Voice Conversion","date":"2019-04-09","arxiv_id":"1904.04631","repositories_listed":6,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/statistical-parametric-speech-synthesis-1","title":"Statistical Parametric Speech Synthesis Incorporating Generative Adversarial Networks","date":"2017-09-23","arxiv_id":"1709.08041","repositories_listed":5,"syntology":null},{"url":"/paper/voice-conversion-from-non-parallel-corpora","title":"Voice Conversion from Non-parallel Corpora Using Variational Auto-encoder","date":"2016-10-13","arxiv_id":"1610.04019","repositories_listed":5,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}},{"url":"/paper/diffusion-based-voice-conversion-with-fast","title":"Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme","date":"2021-09-28","arxiv_id":"2109.13821","repositories_listed":4,"syntology":{"n":9,"n_ran":6,"n_unverified":3,"n_pointer_only":9}},{"url":"/paper/s2vc-a-framework-for-any-to-any-voice","title":"S2VC: A Framework for Any-to-Any Voice Conversion with Self-Supervised Pretrained Representations","date":"2021-04-07","arxiv_id":"2104.02901","repositories_listed":4,"syntology":null},{"url":"/paper/arabert-transformer-based-model-for-arabic","title":"AraBERT: Transformer-based Model for Arabic Language Understanding","date":"2020-02-28","arxiv_id":"2003.00104","repositories_listed":4,"syntology":null},{"url":"/paper/yourtts-towards-zero-shot-multi-speaker-tts","title":"YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone","date":"2021-12-04","arxiv_id":"2112.02418","repositories_listed":3,"syntology":null},{"url":"/paper/maskcyclegan-vc-learning-non-parallel-voice","title":"MaskCycleGAN-VC: Learning Non-parallel Voice Conversion with Filling in Frames","date":"2021-02-25","arxiv_id":"2102.12841","repositories_listed":3,"syntology":null},{"url":"/paper/one-class-learning-towards-generalized-voice-1","title":"One-class learning towards generalized voice spoofing detection","date":"2020-10-27","arxiv_id":"2010.13995","repositories_listed":3,"syntology":{"n":10,"n_ran":0,"n_unverified":10,"n_pointer_only":0}},{"url":"/paper/the-sequence-to-sequence-baseline-for-the","title":"The Sequence-to-Sequence Baseline for the Voice Conversion Challenge 2020: Cascading ASR and TTS","date":"2020-10-06","arxiv_id":"2010.02434","repositories_listed":3,"syntology":null},{"url":"/paper/cotatron-transcription-guided-speech-encoder","title":"Cotatron: Transcription-Guided Speech Encoder for Any-to-Many Voice Conversion without Parallel Data","date":"2020-05-07","arxiv_id":"2005.03295","repositories_listed":3,"syntology":null},{"url":"/paper/stargan-vc2-rethinking-conditional-methods","title":"StarGAN-VC2: Rethinking Conditional Methods for StarGAN-Based Voice Conversion","date":"2019-07-29","arxiv_id":"1907.12279","repositories_listed":3,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":3}},{"url":"/paper/190600794","title":"Blow: a single-scale hyperconditioned flow for non-parallel raw-audio voice conversion","date":"2019-06-03","arxiv_id":"1906.00794","repositories_listed":3,"syntology":{"n":5,"n_ran":1,"n_unverified":4,"n_pointer_only":0}},{"url":"/paper/multi-target-voice-conversion-without","title":"Multi-target Voice Conversion without Parallel Data by Adversarially Learning Disentangled Audio Representations","date":"2018-04-09","arxiv_id":"1804.02812","repositories_listed":3,"syntology":{"n":2,"n_ran":0,"n_unverified":2,"n_pointer_only":2}},{"url":"/paper/audio-deepfake-detection-with-self-supervised-1","title":"Audio Deepfake Detection with Self-Supervised XLS-R and SLS Classifier","date":"2024-10-28","arxiv_id":null,"repositories_listed":2,"syntology":null},{"url":"/paper/hierspeech-bridging-the-gap-between-semantic","title":"HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis","date":"2023-11-21","arxiv_id":"2311.12454","repositories_listed":2,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":1}},{"url":"/paper/evaluating-methods-for-ground-truth-free","title":"Evaluating Methods for Ground-Truth-Free Foreign Accent Conversion","date":"2023-09-05","arxiv_id":"2309.02133","repositories_listed":2,"syntology":null},{"url":"/paper/phoneme-hallucinator-one-shot-voice","title":"Phoneme Hallucinator: One-shot Voice Conversion via Set Expansion","date":"2023-08-11","arxiv_id":"2308.06382","repositories_listed":2,"syntology":null},{"url":"/paper/anonymizing-speech-evaluating-and-designing","title":"Anonymizing Speech: Evaluating and Designing Speaker Anonymization Techniques","date":"2023-08-05","arxiv_id":"2308.04455","repositories_listed":2,"syntology":null},{"url":"/paper/speechlmscore-evaluating-speech-generation","title":"SpeechLMScore: Evaluating speech generation using speech language model","date":"2022-12-08","arxiv_id":"2212.04559","repositories_listed":2,"syntology":{"n":9,"n_ran":1,"n_unverified":8,"n_pointer_only":0}},{"url":"/paper/istftnet-fast-and-lightweight-mel-spectrogram","title":"iSTFTNet: Fast and Lightweight Mel-Spectrogram Vocoder Incorporating Inverse Short-Time Fourier Transform","date":"2022-03-04","arxiv_id":"2203.02395","repositories_listed":2,"syntology":{"n":11,"n_ran":8,"n_unverified":3,"n_pointer_only":0}}],"syntology_records":17,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}