{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-building-text-to-speech-systems-for","title":"Towards Building Text-To-Speech Systems for the Next Billion Users","arxiv_id":"2211.09536","date":"2022-11-17","proceeding":null,"authors":["Gokul Karthik Kumar","Praveen S V","Pratyush Kumar","Mitesh M. Khapra","Karthik Nandakumar"],"abstract":"Deep learning based text-to-speech (TTS) systems have been evolving rapidly with advances in model architectures, training methodologies, and generalization across speakers and languages. However, these advances have not been thoroughly investigated for Indian language speech synthesis. Such investigation is computationally expensive given the number and diversity of Indian languages, relatively lower resource availability, and the diverse set of advances in neural TTS that remain untested. In this paper, we evaluate the choice of acoustic models, vocoders, supplementary loss functions, training schedules, and speaker and language diversity for Dravidian and Indo-Aryan languages. Based on this, we identify monolingual models with FastPitch and HiFi-GAN V1, trained jointly on male and female speakers to perform the best. With this setup, we train and evaluate TTS models for 13 languages and find our models to significantly improve upon existing models in all languages as measured by mean opinion scores. We open-source all models on the Bhashini platform.","url_abs":"https://arxiv.org/abs/2211.09536v3","url_pdf":"https://arxiv.org/pdf/2211.09536v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-building-text-to-speech-systems-for","repo_url":"https://github.com/gokulkarthik/text2speech","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"towards-building-text-to-speech-systems-for","repo_url":"https://github.com/ai4bharat/indic-tts","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"speech-synthesis","task_name":"Speech Synthesis"},{"task_slug":"speech-synthesis-assamese","task_name":"Speech Synthesis - Assamese"},{"task_slug":"speech-synthesis-bengali","task_name":"Speech Synthesis - Bengali"},{"task_slug":"speech-synthesis-bodo","task_name":"Speech Synthesis - Bodo"},{"task_slug":"speech-synthesis-gujarati","task_name":"Speech Synthesis - Gujarati"},{"task_slug":"speech-synthesis-hindi","task_name":"Speech Synthesis - Hindi"},{"task_slug":"speech-synthesis-kannada","task_name":"Speech Synthesis - Kannada"},{"task_slug":"speech-synthesis-malayalam","task_name":"Speech Synthesis - Malayalam"},{"task_slug":"speech-synthesis-manipuri","task_name":"Speech Synthesis - Manipuri"},{"task_slug":"speech-synthesis-marathi","task_name":"Speech Synthesis - Marathi"},{"task_slug":"speech-synthesis-rajasthani","task_name":"Speech Synthesis - Rajasthani"},{"task_slug":"speech-synthesis-tamil","task_name":"Speech Synthesis - Tamil"},{"task_slug":"speech-synthesis-telugu","task_name":"Speech Synthesis - Telugu"},{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"text-to-speech-synthesis","task_name":"Text-To-Speech Synthesis"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"}],"methods":[{"method_slug":"attention","method_name":"Attention"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"fastpitch","method_name":"FastPitch"},{"method_slug":"hifi-gan","method_name":"HiFi-GAN"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-synthesis-assamese-on-indictts","task":"Speech Synthesis - Assamese","dataset":"IndicTTS","model":"AI4BharatTTS - FastPitch with HiFiGAN","rank_in_archive_order":1,"of":1,"metrics":{"Mean Opinion Score":"2.39"},"uses_additional_data":false},{"leaderboard":"/sota/speech-synthesis-bengali-on-indictts","task":"Speech Synthesis - Bengali","dataset":"IndicTTS","model":"AI4BharatTTS - FastPitch with HiFiGAN","rank_in_archive_order":1,"of":1,"metrics":{"Mean Opinion Score":"3.37"},"uses_additional_data":false},{"leaderboard":"/sota/speech-synthesis-bodo-on-indictts","task":"Speech Synthesis - Bodo","dataset":"IndicTTS","model":"AI4BharatTTS - FastPitch with HiFiGAN","rank_in_archive_order":1,"of":1,"metrics":{"Mean Opinion Score":"3.53"},"uses_additional_data":false},{"leaderboard":"/sota/speech-synthesis-gujarati-on-indictts","task":"Speech Synthesis - Gujarati","dataset":"IndicTTS","model":"AI4BharatTTS - FastPitch with HiFiGAN","rank_in_archive_order":1,"of":1,"metrics":{"Mean Opinion Score":"3.58"},"uses_additional_data":false},{"leaderboard":"/sota/speech-synthesis-hindi-on-indictts","task":"Speech Synthesis - Hindi","dataset":"IndicTTS","model":"AI4BharatTTS - FastPitch with HiFiGAN","rank_in_archive_order":1,"of":1,"metrics":{"Mean Opinion Score":"4.00"},"uses_additional_data":false},{"leaderboard":"/sota/speech-synthesis-kannada-on-indictts","task":"Speech Synthesis - Kannada","dataset":"IndicTTS","model":"AI4BharatTTS - FastPitch with HiFiGAN","rank_in_archive_order":1,"of":1,"metrics":{"Mean Opinion Score":"3.68"},"uses_additional_data":false},{"leaderboard":"/sota/speech-synthesis-malayalam-on-indictts","task":"Speech Synthesis - Malayalam","dataset":"IndicTTS","model":"AI4BharatTTS - FastPitch with HiFiGAN","rank_in_archive_order":1,"of":1,"metrics":{"Mean Opinion Score":"3.64"},"uses_additional_data":false},{"leaderboard":"/sota/speech-synthesis-manipuri-on-indictts","task":"Speech Synthesis - Manipuri","dataset":"IndicTTS","model":"AI4BharatTTS - FastPitch with HiFiGAN","rank_in_archive_order":1,"of":1,"metrics":{"Mean Opinion Score":"3.30"},"uses_additional_data":false},{"leaderboard":"/sota/speech-synthesis-marathi-on-indictts","task":"Speech Synthesis - Marathi","dataset":"IndicTTS","model":"AI4BharatTTS - FastPitch with HiFiGAN","rank_in_archive_order":1,"of":1,"metrics":{"Mean Opinion Score":"3.26"},"uses_additional_data":false},{"leaderboard":"/sota/speech-synthesis-rajasthani-on-indictts","task":"Speech Synthesis - Rajasthani","dataset":"IndicTTS","model":"AI4BharatTTS - FastPitch with HiFiGAN","rank_in_archive_order":1,"of":1,"metrics":{"Mean Opinion Score":"3.40"},"uses_additional_data":false},{"leaderboard":"/sota/speech-synthesis-tamil-on-indictts","task":"Speech Synthesis - Tamil","dataset":"IndicTTS","model":"AI4BharatTTS - FastPitch with HiFiGAN","rank_in_archive_order":1,"of":1,"metrics":{"Mean Opinion Score":"3.84"},"uses_additional_data":false},{"leaderboard":"/sota/speech-synthesis-telugu-on-indictts","task":"Speech Synthesis - Telugu","dataset":"IndicTTS","model":"AI4BharatTTS - FastPitch with HiFiGAN","rank_in_archive_order":1,"of":1,"metrics":{"Mean Opinion Score":"3.66"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2211.09536","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2211.09536"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ai4bharat/indic-tts","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/gokulkarthik/text2speech","reach":{"status":"ok"}}],"summary":{"unverified":1},"by_repo_kind":{"listed":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"d8144ca77e15f25f","entry":"formatter_indictts","repo":"ai4bharat/indic-tts","repo_kind":"listed","path":"vocoder.py","file_url":"https://github.com/ai4bharat/indic-tts/blob/HEAD/vocoder.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d8144ca77e15f25f"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}