{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/190501965","title":"Arabic Text Diacritization Using Deep Neural Networks","arxiv_id":"1905.01965","date":"2019-04-25","proceeding":null,"authors":["Ali Fadel","Ibraheem Tuffaha","Bara' Al-Jawarneh","Mahmoud Al-Ayyoub"],"abstract":"Diacritization of Arabic text is both an interesting and a challenging\nproblem at the same time with various applications ranging from speech\nsynthesis to helping students learning the Arabic language. Like many other\ntasks or problems in Arabic language processing, the weak efforts invested into\nthis problem and the lack of available (open-source) resources hinder the\nprogress towards solving this problem. This work provides a critical review for\nthe currently existing systems, measures and resources for Arabic text\ndiacritization. Moreover, it introduces a much-needed free-for-all cleaned\ndataset that can be easily used to benchmark any work on Arabic diacritization.\nExtracted from the Tashkeela Corpus, the dataset consists of 55K lines\ncontaining about 2.3M words. After constructing the dataset, existing tools and\nsystems are tested on it. The results of the experiments show that the neural\nShakkala system significantly outperforms traditional rule-based approaches and\nother closed-source tools with a Diacritic Error Rate (DER) of 2.88% compared\nwith 13.78%, which the best DER for the non-neural approach (obtained by the\nMishkal tool).","url_abs":"http://arxiv.org/abs/1905.01965v1","url_pdf":"http://arxiv.org/pdf/1905.01965v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"190501965","repo_url":"https://github.com/AliOsm/arabic-text-diacritization","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"190501965","repo_url":"https://github.com/Barqawiz/Shakkala","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"arabic-text-diacritization","task_name":"Arabic Text Diacritization"}],"methods":[],"datasets_introduced":[{"slug":"arabic-text-diacritization","name":"Arabic Text Diacritization","full_name":null}],"methods_introduced":[],"results":[{"leaderboard":"/sota/arabic-text-diacritization-on-tashkeela-1","task":"Arabic Text Diacritization","dataset":"Tashkeela","model":"Shakkala","rank_in_archive_order":6,"of":6,"metrics":{"Diacritic Error Rate":"0.0373","Word Error Rate (WER)":"0.1119"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}