Browse State-of-the-Art › Arabic Text Diacritization
Arabic Text Diacritization
7 papers with code · 2 benchmarks · 3 datasets archive 2025-07-28
Addition of diacritics for undiacritized arabic texts for words disambiguation.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CATT (12 rows) | CATT ED | CATT: Character-based Arabic Tashkeel Transformer | code | — | Compare |
| Tashkeela (6 rows) | CBHG model | Effective Deep Learning Models for Automatic Diacritization of Arabic Text | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
7 shown of 7 papers with code (13 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
8 Nov 2019 2 repositories listedIn this work, we present several deep learning models for the automatic diacritization of Arabic text.
-
25 Apr 2019 2 repositories listedAfter constructing the dataset, existing tools and systems are tested on it.
-
3 Jul 2024 1 repository listedThen, we applied the Noisy-Student approach to boost the performance of the best model.
-
1 Nov 2020 1 repository listedWe propose a novel architecture for labelling character sequences that achieves state-of-the-art results on the Tashkeela Arabic diacritization benchmark.
-
1 Nov 2020 1 repository listedWe propose three deep learning models to recover Arabic text diacritics based on our work in a text-to-speech synthesis system using deep learning.
-
1 May 2020 1 repository listedWe present CAMeL Tools, a collection of open-source tools for Arabic natural language processing in Python.
-
8 Apr 2020 1 repository listedIn this paper, we propose an approach to tackle the problem of the automatic restoration of Arabic diacritics that includes three components stacked in a pipeline: a deep learning model which is a multi-layer recurrent…
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections