Browse State-of-the-Art › Voice Conversion › Papers, page 3
Voice Conversion
Papers archive 2025-07-28
archive papers tagged: 520 · with a code link: 175 · where Syntology ran a sample: 41 (32 with a run with no instrument failure, 9 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (41 of 520 tagged: 32 with a run with no instrument failure, 9 where every run was a failure of Syntology's instrument)
Page 3 of 6: papers 201 to 300 of 520, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Mitigating Timbre Leakage with Universal Semantic Mapping Residual Block for Voice Conversion11 Apr 2025 0 repositories listed
-
An Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASR11 Mar 2025 0 repositories listed
-
ASVspoof 5: Design, Collection and Validation of Resources for Spoofing, Deepfake, and Adversarial Attack Detection Using Crowdsourced Speech13 Feb 2025 0 repositories listed
-
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement11 Feb 2025 0 repositories listed
-
Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody Features7 Feb 2025 0 repositories listed
-
FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks6 Feb 2025 0 repositories listed
-
GenVC: Self-Supervised Zero-Shot Voice Conversion6 Feb 2025 0 repositories listed
-
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching29 Jan 2025 0 repositories listed
-
Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning26 Jan 2025 0 repositories listed
-
Generalizable Audio Deepfake Detection via Latent Space Refinement and Augmentation24 Jan 2025 0 repositories listed
-
Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR17 Jan 2025 0 repositories listed
-
Speech Synthesis along Perceptual Voice Quality Dimensions15 Jan 2025 0 repositories listed
-
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives11 Jan 2025 0 repositories listed
-
ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training8 Jan 2025 0 repositories listed
-
Generating and Detecting Various Types of Fake Image and Audio Content: A Review of Modern Deep Learning Technologies and Tools7 Jan 2025 0 repositories listed
-
AdaptVC: High Quality Voice Conversion with Adaptive Learning2 Jan 2025 0 repositories listed
-
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion29 Dec 2024 0 repositories listed
-
A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction11 Dec 2024 0 repositories listed
-
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching6 Dec 2024 0 repositories listed
-
Noro: A Noise-Robust One-shot Voice Conversion System with Hidden Speaker Representation Capabilities29 Nov 2024 0 repositories listed
-
SKQVC: One-Shot Voice Conversion by K-Means Quantization with Self-Supervised Speech Representations25 Nov 2024 0 repositories listed
-
CTEFM-VC: Zero-Shot Voice Conversion Based on Content-Aware Timbre Ensemble Modeling and Flow Matching4 Nov 2024 0 repositories listed
-
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec21 Oct 2024 0 repositories listed
-
Improving Voice Quality in Speech Anonymization With Just Perception-Informed Losses20 Oct 2024 0 repositories listed
-
A Pilot Study of Applying Sequence-to-Sequence Voice Conversion to Evaluate the Intelligibility of L2 Speech Using a Native Speaker's Shadowings3 Oct 2024 0 repositories listed
-
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling2 Oct 2024 0 repositories listed
-
Exploring synthetic data for cross-speaker style transfer in style representation based TTS25 Sep 2024 0 repositories listed
-
Textless NLP -- Zero Resource Challenge with Low Resource Compute24 Sep 2024 0 repositories listed
-
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion17 Sep 2024 0 repositories listed
-
HLTCOE JHU Submission to the Voice Privacy Challenge 202413 Sep 2024 0 repositories listed
-
LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling13 Sep 2024 0 repositories listed
-
D-CAPTCHA++: A Study of Resilience of Deepfake CAPTCHA under Transferable Imperceptible Adversarial Attack11 Sep 2024 0 repositories listed
-
VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion10 Sep 2024 0 repositories listed
-
VoiceWukong: Benchmarking Deepfake Voice Detection10 Sep 2024 0 repositories listed
-
ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism5 Sep 2024 0 repositories listed
-
Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder5 Sep 2024 0 repositories listed
-
FastVoiceGrad: One-step Diffusion-Based Voice Conversion with Adversarial Conditional Diffusion Distillation3 Sep 2024 0 repositories listed
-
Pureformer-VC: Non-parallel One-Shot Voice Conversion with Pure Transformer Blocks and Triplet Discriminative Training3 Sep 2024 0 repositories listed
-
USTC-KXDIGIT System Description for ASVspoof5 Challenge3 Sep 2024 0 repositories listed
-
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders3 Sep 2024 0 repositories listed
-
Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion1 Sep 2024 0 repositories listed
-
Progressive Residual Extraction based Pre-training for Speech Representation Learning31 Aug 2024 0 repositories listed
-
AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge30 Aug 2024 0 repositories listed
-
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models28 Aug 2024 0 repositories listed
-
MaskCycleGAN-based Whisper to Normal Speech Conversion27 Aug 2024 0 repositories listed
-
Toward Improving Synthetic Audio Spoofing Detection Robustness via Meta-Learning and Disentangled Training With Adversarial Examples23 Aug 2024 0 repositories listed
-
LCM-SVC: Latent Diffusion Model Based Singing Voice Conversion with Inference Acceleration via Latent Consistency Distillation22 Aug 2024 0 repositories listed
-
VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing11 Aug 2024 0 repositories listed
-
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency8 Aug 2024 0 repositories listed
-
Automatic Voice Identification after Speech Resynthesis using PPG5 Aug 2024 0 repositories listed
-
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion5 Aug 2024 0 repositories listed
-
Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity20 Jul 2024 0 repositories listed
-
The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice Anonymisation16 Jul 2024 0 repositories listed
-
Source Tracing of Audio Deepfake Systems10 Jul 2024 0 repositories listed
-
We Need Variations in Speech Generation: Sub-center Modelling for Speaker Embeddings5 Jul 2024 0 repositories listed
-
Application of ASV for Voice Identification after VC and Duration Predictor Improvement in TTS Models27 Jun 2024 0 repositories listed
-
24 Jun 2024 0 repositories listed
-
RefXVC: Cross-Lingual Voice Conversion with Enhanced Reference Leveraging24 Jun 2024 0 repositories listed
-
DualVC 3: Leveraging Language Model Generated Pseudo Context for End-to-end Low Latency Streaming Voice Conversion12 Jun 2024 0 repositories listed
-
Improving child speech recognition with augmented child-like speech12 Jun 2024 0 repositories listed
-
SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models12 Jun 2024 0 repositories listed
-
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion9 Jun 2024 0 repositories listed
-
LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance8 Jun 2024 0 repositories listed
-
The Database and Benchmark for the Source Speaker Tracing Challenge 20247 Jun 2024 0 repositories listed
-
Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline6 Jun 2024 0 repositories listed
-
Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion4 Jun 2024 0 repositories listed
-
Real-Time and Accurate: Zero-shot High-Fidelity Singing Voice Conversion with Multi-Condition Flow Synthesis23 May 2024 0 repositories listed
-
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model2 May 2024 0 repositories listed
-
Who is Authentic Speaker30 Apr 2024 0 repositories listed
-
Retrieval-Augmented Audio Deepfake Detection22 Apr 2024 0 repositories listed
-
PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders3 Apr 2024 0 repositories listed
-
Voice Conversion Augmentation for Speaker Recognition on Defective Datasets1 Apr 2024 0 repositories listed
-
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion3 Mar 2024 0 repositories listed
-
Transcription and translation of videos using fine-tuned XLSR Wav2Vec2 on custom dataset and mBART1 Mar 2024 0 repositories listed
-
Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations5 Feb 2024 0 repositories listed
-
SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition31 Jan 2024 0 repositories listed
-
SongBsAb: A Dual Prevention Approach against Singing Voice Conversion based Illegal Song Covers30 Jan 2024 0 repositories listed
-
Adversarial speech for voice privacy protection from Personalized Speech generation22 Jan 2024 0 repositories listed
-
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion19 Jan 2024 0 repositories listed
-
Transfer the linguistic representations from TTS to accent conversion with non-parallel data7 Jan 2024 0 repositories listed
-
StreamVC: Real-Time Low-Latency Voice Conversion5 Jan 2024 0 repositories listed
-
Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion29 Dec 2023 0 repositories listed
-
AE-Flow: AutoEncoder Normalizing Flow27 Dec 2023 0 repositories listed
-
Exploring data augmentation in bias mitigation against non-native-accented speech24 Dec 2023 0 repositories listed
-
Creating New Voices using Normalizing Flows22 Dec 2023 0 repositories listed
-
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention14 Dec 2023 0 repositories listed
-
PerMod: Perceptually Grounded Voice Modification with Latent Diffusion Models13 Dec 2023 0 repositories listed
-
Vulnerability of Automatic Identity Recognition to Audio-Visual Deepfakes29 Nov 2023 0 repositories listed
-
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion24 Nov 2023 0 repositories listed
-
Reimagining Speech: A Scoping Review of Deep Learning-Powered Voice Conversion14 Nov 2023 0 repositories listed
-
Parrot-Trained Adversarial Examples: Pushing the Practicality of Black-Box Audio Attacks against Speaker Recognition Models13 Nov 2023 0 repositories listed
-
An overview of text-to-speech systems and media applications22 Oct 2023 0 repositories listed
-
SelfVC: Voice Conversion With Iterative Refinement using Self Transformations14 Oct 2023 0 repositories listed
-
Voice Conversion for Stuttered Speech, Instruments, Unseen Languages and Textually Described Voices12 Oct 2023 0 repositories listed
-
AutoCycle-VC: Towards Bottleneck-Independent Zero-Shot Cross-Lingual Voice Conversion10 Oct 2023 0 repositories listed
-
A Comparative Study of Voice Conversion Models with Large-Scale Speech and Singing Data: The T13 Systems for the Singing Voice Conversion Challenge 20238 Oct 2023 0 repositories listed
-
VITS-Based Singing Voice Conversion Leveraging Whisper and multi-scale F0 Modeling4 Oct 2023 0 repositories listed
-
DualVC 2: Dynamic Masked Convolution for Unified Streaming and Non-Streaming Voice Conversion27 Sep 2023 0 repositories listed
-
Towards General-Purpose Text-Instruction-Guided Voice Conversion25 Sep 2023 0 repositories listed
-
The Impact of Silence on Speech Anti-Spoofing21 Sep 2023 0 repositories listed