{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/attentive-sequence-to-sequence-learning-for","title":"Attentive Sequence-to-Sequence Learning for Diacritic Restoration of Yorùbá Language Text","arxiv_id":"1804.00832","date":"2018-04-03","proceeding":null,"authors":["Iroro Orife"],"abstract":"Yor\\`ub\\'a is a widely spoken West African language with a writing system\nrich in tonal and orthographic diacritics. With very few exceptions, diacritics\nare omitted from electronic texts, due to limited device and application\nsupport. Diacritics provide morphological information, are crucial for lexical\ndisambiguation, pronunciation and are vital for any Yor\\`ub\\'a text-to-speech\n(TTS), automatic speech recognition (ASR) and natural language processing (NLP)\ntasks. Reframing Automatic Diacritic Restoration (ADR) as a machine translation\ntask, we experiment with two different attentive Sequence-to-Sequence neural\nmodels to process undiacritized text. On our evaluation dataset, this approach\nproduces diacritization error rates of less than 5%. We have released\npre-trained models, datasets and source-code as an open-source project to\nadvance efforts on Yor\\`ub\\'a language technology.","url_abs":"http://arxiv.org/abs/1804.00832v2","url_pdf":"http://arxiv.org/pdf/1804.00832v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"attentive-sequence-to-sequence-learning-for","repo_url":"https://github.com/Niger-Volta-LTI/yoruba-adr","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"automatic-speech-recognition-2","task_name":"Automatic Speech Recognition"},{"task_slug":"automatic-speech-recognition","task_name":"Automatic Speech Recognition (ASR)"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.00832","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}