{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-very-low-resource-language-speech-corpus","title":"A Very Low Resource Language Speech Corpus for Computational Language Documentation Experiments","arxiv_id":"1710.03501","date":"2017-10-10","proceeding":"LREC 2018 5","authors":["P. Godard","G. Adda","M. Adda-Decker","J. Benjumea","L. Besacier","J. Cooper-Leavitt","G-N. Kouarata","L. Lamel","H. Maynard","M. Mueller","A. Rialland","S. Stueker","F. Yvon","M. Zanon-Boito"],"abstract":"Most speech and language technologies are trained with massive amounts of\nspeech and text information. However, most of the world languages do not have\nsuch resources or stable orthography. Systems constructed under these almost\nzero resource conditions are not only promising for speech technology but also\nfor computational language documentation. The goal of computational language\ndocumentation is to help field linguists to (semi-)automatically analyze and\nannotate audio recordings of endangered and unwritten languages. Example tasks\nare automatic phoneme discovery or lexicon discovery from the speech signal.\nThis paper presents a speech corpus collected during a realistic language\ndocumentation process. It is made up of 5k speech utterances in Mboshi (Bantu\nC25) aligned to French text translations. Speech transcriptions are also made\navailable: they correspond to a non-standard graphemic form close to the\nlanguage phonology. We present how the data was collected, cleaned and\nprocessed and we illustrate its use through a zero-resource task: spoken term\ndiscovery. The dataset is made available to the community for reproducible\ncomputational language documentation experiments and their evaluation.","url_abs":"http://arxiv.org/abs/1710.03501v3","url_pdf":"http://arxiv.org/pdf/1710.03501v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-very-low-resource-language-speech-corpus","repo_url":"https://github.com/besacier/mboshi-french-parallel-corpus","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"a-very-low-resource-language-speech-corpus","repo_url":"https://github.com/mzboito/mmboshi","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1710.03501","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}