{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bibletts-a-large-high-fidelity-multilingual","title":"BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus","arxiv_id":"2207.03546","date":"2022-07-07","proceeding":null,"authors":["Josh Meyer","David Ifeoluwa Adelani","Edresson Casanova","Alp Öktem","Daniel Whitenack Julian Weber","Salomon Kabongo","Elizabeth Salesky","Iroro Orife","Colin Leong","Perez Ogayo","Chris Emezue","Jonathan Mukiibi","Salomey Osei","Apelete Agbolo","Victor Akinode","Bernard Opoku","Samuel Olanrewaju","Jesujoba Alabi","Shamsuddeen Muhammad"],"abstract":"BibleTTS is a large, high-quality, open speech dataset for ten languages spoken in Sub-Saharan Africa. The corpus contains up to 86 hours of aligned, studio quality 48kHz single speaker recordings per language, enabling the development of high-quality text-to-speech models. The ten languages represented are: Akuapem Twi, Asante Twi, Chichewa, Ewe, Hausa, Kikuyu, Lingala, Luganda, Luo, and Yoruba. This corpus is a derivative work of Bible recordings made and released by the Open.Bible project from Biblica. We have aligned, cleaned, and filtered the original recordings, and additionally hand-checked a subset of the alignments for each language. We present results for text-to-speech models with Coqui TTS. The data is released under a commercial-friendly CC-BY-SA license.","url_abs":"https://arxiv.org/abs/2207.03546v1","url_pdf":"https://arxiv.org/pdf/2207.03546v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bibletts-a-large-high-fidelity-multilingual","repo_url":"https://github.com/alpoktem/bible2speechdb","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2207.03546","atlas_url":"https://app.syntology.ai/?focus=2207.03546","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}