Browse State-of-the-Art › Voice Cloning
Voice Cloning
36 papers with code · 0 benchmarks · 2 datasets archive 2025-07-28
Voice cloning is a highly desired feature for personalized speech interfaces. Neural voice cloning system learns to synthesize a person’s voice from only a few audio samples.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 36 papers with code (112 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
12 Jun 2018 11 repositories listedClone a voice in 5 seconds to generate arbitrary speech in real-time
-
9 Jul 2019 4 repositories listed Syntology ran 2 of 9 samples · 7 unverified · 7 pointer-only (licence)We present a multispeaker, multilingual text-to-speech (TTS) synthesis model based on Tacotron that is able to produce high quality speech in multiple languages.
-
4 Jul 2024 3 repositories listed Syntology ran 10 of 10 samples · 0 unverifiedThis report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs).
-
18 Apr 2023 3 repositories listedKeywords: Bark, ai voice cloning, Suno, text-to-speech, artificial intelligence, audio generation, Meta's encodec, audio codebooks, semantic tokens, HuBert, transformer-based model, multilingual speech, wav2vec, linear…
-
14 Aug 2024 2 repositories listedAudio deepfake detection has become a pivotal task over the last couple of years, as many recent speech synthesis and voice cloning systems generate highly realistic speech samples, thus enabling their use in malicious…
-
5 Aug 2023 2 repositories listedThe growing use of voice user interfaces has led to a surge in the collection and storage of speech data.
-
7 Nov 2022 2 repositories listedIn this paper, we extend the pretraining method for cross-lingual multi-speaker speech synthesis tasks, including cross-lingual multi-speaker voice cloning and cross-lingual multi-speaker speech editing.
-
14 Feb 2018 2 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedSpeaker adaptation is based on fine-tuning a multi-speaker generative model with a few cloning samples.
-
29 May 2025 1 repository listedRecent advancements in Text-to-Speech (TTS) models, particularly in voice cloning, have intensified the demand for adaptable and efficient deepfake detection methods.
-
29 Apr 2025 1 repository listedWe present a novel benchmark for voice cloning text-to-speech models.
-
31 Mar 2025 1 repository listedHigh-quality speech dialogue datasets are crucial for Speech-LLM development, yet existing acquisition methods face significant limitations.
-
3 Mar 2025 1 repository listed Syntology ran 1 of 2 samples · 1 unverifiedRecent advancements in large language models (LLMs) have driven significant progress in zero-shot text-to-speech (TTS) synthesis.
-
18 Feb 2025 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedText-to-song generation, the task of creating vocals and accompaniment from textual inputs, poses significant challenges due to domain complexity and data scarcity.
-
17 Feb 2025 1 repository listedBased on our new StepEval-Audio-360 evaluation benchmark, Step-Audio achieves state-of-the-art performance in human evaluations, especially in terms of instruction following.
-
8 Feb 2025 1 repository listedRecently, large language model (LLM) based text-to-speech (TTS) systems have gradually become the mainstream in the industry due to their high naturalness and powerful zero-shot voice cloning capabilities.
-
30 Oct 2024 1 repository listedHowever, this approach is limited in its ability to handle numerous or lengthy speech excerpts, since the concatenation of source and target speech must fall within the maximum context length which is determined during…
-
1 Oct 2024 1 repository listedWhile recent advances in Text-to-Speech (TTS) technology produce natural and expressive speech, they lack the option for users to select emotion and control intensity.
-
23 Sep 2024 1 repository listedPrevious fake speech datasets were constructed from a defender's perspective to develop countermeasure (CM) systems without considering diverse motivations of attackers.
-
27 Aug 2024 1 repository listedSeven SOTA audio spoof detection approaches are evaluated on this laundered database.
-
7 Jul 2024 1 repository listedBased on the tokens, we further propose a scalable zero-shot TTS synthesizer, CosyVoice, which consists of an LLM for text-to-token generation and a conditional flow matching model for token-to-speech synthesis.
-
7 Jun 2024 1 repository listed Syntology ran 3 of 6 samples · 3 unverifiedMost Zero-shot Multi-speaker TTS (ZS-TTS) systems support only a single language.
-
6 Jun 2024 1 repository listedRecent advancements in text-to-speech (TTS) powered by language models have showcased remarkable capabilities in achieving naturalness and zero-shot voice cloning.
-
14 May 2024 1 repository listedWe conduct comprehensive experiments using state-of-the-art detection methods on PolyGlotFake dataset.
-
20 Feb 2024 1 repository listedGiven a script, the challenge in Movie Dubbing (Visual Voice Cloning, V2C) is to generate speech that aligns well with the video in both time and emotion, based on the tone of a reference audio track.
-
30 Jan 2024 1 repository listed Syntology ran 1 of 3 samples · 2 unverifiedIn the rapidly evolving field of speech generative models, there is a pressing need to ensure audio authenticity against the risks of voice cloning.
-
3 Dec 2023 1 repository listed Syntology ran 3 of 15 samples · 12 unverifiedThe voice styles are not directly copied from and constrained by the style of the reference speaker.
-
15 Jul 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverifiedSynthetic-voice cloning technologies have seen significant advances in recent years, giving rise to a range of potential harms.
-
21 Oct 2022 1 repository listedWhile neural methods for text-to-speech (TTS) have shown great advances in modeling multiple speakers, even in zero-shot settings, the amount of data needed for those approaches is generally not feasible for the vast…
-
14 Oct 2022 1 repository listedWe present a comprehensive empirical study for personalized spontaneous speech synthesis on the basis of linguistic knowledge.
-
24 Apr 2022 1 repository listedIn this paper, we propose dictionary attacks against speaker verification - a novel attack vector that aims to match a large fraction of speaker population by chance.
Syntology lines on 9 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections