{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/wanna-hear-your-voice-adaptive-effective-and","title":"Wanna hear your voice? A sample is all we need!","arxiv_id":"2410.00527","date":"2024-10-01","proceeding":null,"authors":["The Hieu Pham","Phuong Thanh Tran Nguyen","Xuan Tho Nguyen","Tan Dat Nguyen","Duc Dung Nguyen"],"abstract":"Research on audio clue-based target speaker extraction (TSE) has focused on modeling mixtures and reference speech, achieving strong results in English due to abundant datasets. However, cross-lingual properties remain underexplored, as low-resource languages face challenges from limited annotated data and linguistic resources. To bridge this gap, we propose WHYV (Wanna Hear Your Voice), a cross-lingual TSE framework enabling zero-shot adaptation without fine-tuning. WHYV employs a frequency-modulated gating mechanism that dynamically adjusts the acoustic features of the target speaker, minimizing reliance on language-specific cues. Evaluations demonstrate state-of-the-art zero-shot performance: 13.8 dB (Libri2Mix mix-both), 18.1 dB (mix-clean), and 14.8 dB on Vietnamese data.","url_abs":"https://arxiv.org/abs/2410.00527v4","url_pdf":"https://arxiv.org/pdf/2410.00527v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"all","task_name":"All"},{"task_slug":"speech-separation","task_name":"Speech Separation"},{"task_slug":"target-speaker-extraction","task_name":"Target Speaker Extraction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-separation-on-libri2mix","task":"Speech Separation","dataset":"Libri2Mix","model":"WHYV","rank_in_archive_order":5,"of":10,"metrics":{"SDR":"17.2458","SI-SDRi":"17.5"},"uses_additional_data":false},{"leaderboard":"/sota/speech-separation-on-wham","task":"Speech Separation","dataset":"WHAM!","model":"WHYV","rank_in_archive_order":6,"of":6,"metrics":{"SI-SDRi":"12.964"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}