{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/indicvoices-r-unlocking-a-massive","title":"IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS","arxiv_id":"2409.05356","date":"2024-09-09","proceeding":null,"authors":["Ashwin Sankar","Srija Anand","Praveen Srinivasa Varadhan","Sherry Thomas","Mehak Singal","Shridhar Kumar","Deovrat Mehendale","Aditi Krishana","Giri Raju","Mitesh Khapra"],"abstract":"Recent advancements in text-to-speech (TTS) synthesis show that large-scale models trained with extensive web data produce highly natural-sounding output. However, such data is scarce for Indian languages due to the lack of high-quality, manually subtitled data on platforms like LibriVox or YouTube. To address this gap, we enhance existing large-scale ASR datasets containing natural conversations collected in low-quality environments to generate high-quality TTS training data. Our pipeline leverages the cross-lingual generalization of denoising and speech enhancement models trained on English and applied to Indian languages. This results in IndicVoices-R (IV-R), the largest multilingual Indian TTS dataset derived from an ASR dataset, with 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. IV-R matches the quality of gold-standard TTS datasets like LJSpeech, LibriTTS, and IndicTTS. We also introduce the IV-R Benchmark, the first to assess zero-shot, few-shot, and many-shot speaker generalization capabilities of TTS models on Indian voices, ensuring diversity in age, gender, and style. We demonstrate that fine-tuning an English pre-trained model on a combined dataset of high-quality IndicTTS and our IV-R dataset results in better zero-shot speaker generalization compared to fine-tuning on the IndicTTS dataset alone. Further, our evaluation reveals limited zero-shot generalization for Indian voices in TTS models trained on prior datasets, which we improve by fine-tuning the model on our data containing diverse set of speakers across language families. We open-source all data and code, releasing the first TTS model for all 22 official Indian languages.","url_abs":"https://arxiv.org/abs/2409.05356v2","url_pdf":"https://arxiv.org/pdf/2409.05356v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"indicvoices-r-unlocking-a-massive","repo_url":"https://github.com/ai4bharat/indicvoices-r","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"CC-BY-4.0"}}],"tasks":[{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"},{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"zero-shot-generalization","task_name":"Zero-shot Generalization"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2409.05356","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2409.05356"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ai4bharat/indicvoices-r","reach":{"status":"ok","spdx":"CC-BY-4.0"}}],"summary":{"ran":10,"unverified":1},"by_repo_kind":{"official":{"samples":11,"ran":10,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":11,"samples":[{"code_sha256_prefix":"3e228ff303bcad94","entry":"contains_english_characters","repo":"ai4bharat/indicvoices-r","repo_kind":"official","path":"VoiceCraft/inference.py","file_url":"https://github.com/ai4bharat/indicvoices-r/blob/HEAD/VoiceCraft/inference.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"3e228ff303bcad94"}},{"code_sha256_prefix":"06244bf64882ebfe","entry":"find_closest_cut_off_word","repo":"ai4bharat/indicvoices-r","repo_kind":"official","path":"VoiceCraft/predict.py","file_url":"https://github.com/ai4bharat/indicvoices-r/blob/HEAD/VoiceCraft/predict.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"06244bf64882ebfe"}},{"code_sha256_prefix":"e0162cfc7ba3b31c","entry":"get_mask_interval_from_word_bounds","repo":"ai4bharat/indicvoices-r","repo_kind":"official","path":"VoiceCraft/predict.py","file_url":"https://github.com/ai4bharat/indicvoices-r/blob/HEAD/VoiceCraft/predict.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"e0162cfc7ba3b31c"}},{"code_sha256_prefix":"ae9e5db4e53c6dce","entry":"get_span","repo":"ai4bharat/indicvoices-r","repo_kind":"official","path":"VoiceCraft/edit_utils.py","file_url":"https://github.com/ai4bharat/indicvoices-r/blob/HEAD/VoiceCraft/edit_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"ae9e5db4e53c6dce"}},{"code_sha256_prefix":"1361930f16042491","entry":"get_spk2item","repo":"ai4bharat/indicvoices-r","repo_kind":"official","path":"VoiceCraft/inference.py","file_url":"https://github.com/ai4bharat/indicvoices-r/blob/HEAD/VoiceCraft/inference.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"1361930f16042491"}},{"code_sha256_prefix":"a6e1062a6dea09b3","entry":"get_transcribe_state","repo":"ai4bharat/indicvoices-r","repo_kind":"official","path":"VoiceCraft/gradio_app.py","file_url":"https://github.com/ai4bharat/indicvoices-r/blob/HEAD/VoiceCraft/gradio_app.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"a6e1062a6dea09b3"}},{"code_sha256_prefix":"76f8432650c37dd1","entry":"get_transcribe_state","repo":"ai4bharat/indicvoices-r","repo_kind":"official","path":"VoiceCraft/predict.py","file_url":"https://github.com/ai4bharat/indicvoices-r/blob/HEAD/VoiceCraft/predict.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"76f8432650c37dd1"}},{"code_sha256_prefix":"c0e30f6224e1afda","entry":"read_jsonl","repo":"ai4bharat/indicvoices-r","repo_kind":"official","path":"VoiceCraft/inference.py","file_url":"https://github.com/ai4bharat/indicvoices-r/blob/HEAD/VoiceCraft/inference.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"c0e30f6224e1afda"}},{"code_sha256_prefix":"99a311d99448f881","entry":"top_k_top_p_filtering","repo":"ai4bharat/indicvoices-r","repo_kind":"official","path":"VoiceCraft/models/voicecraft.py","file_url":"https://github.com/ai4bharat/indicvoices-r/blob/HEAD/VoiceCraft/models/voicecraft.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"99a311d99448f881"}},{"code_sha256_prefix":"fb8e21d446dcf26f","entry":"topk_sampling","repo":"ai4bharat/indicvoices-r","repo_kind":"official","path":"VoiceCraft/models/voicecraft.py","file_url":"https://github.com/ai4bharat/indicvoices-r/blob/HEAD/VoiceCraft/models/voicecraft.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"fb8e21d446dcf26f"}},{"code_sha256_prefix":"d3cb669c6fbc4bf8","entry":"find_closest_word_boundary","repo":"ai4bharat/indicvoices-r","repo_kind":"official","path":"VoiceCraft/tts_demo.py","file_url":"https://github.com/ai4bharat/indicvoices-r/blob/HEAD/VoiceCraft/tts_demo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"d3cb669c6fbc4bf8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}