Papers › Automatic Dialect Detection in Arabic Broadcast Speech
Automatic Dialect Detection in Arabic Broadcast Speech
Ahmed Ali, Najim Dehak, Patrick Cardinal, Sameer Khurana, Sree Harsha Yella, James Glass, Peter Bell, Steve Renals
We investigate different approaches for dialect identification in Arabic broadcast speech, using phonetic, lexical features obtained from a speech recognition system, and acoustic features using the i-vector framework. We studied both generative and discriminate classifiers, and we combined these features using a multi-class Support Vector Machine (SVM). We validated our results on an Arabic/English language identification task, with an accuracy of 100%. We used these features in a binary classifier to discriminate between Modern Standard Arabic (MSA) and Dialectal Arabic, with an accuracy of 100%. We further report results using the proposed method to discriminate between the five most widely used dialects of Arabic: namely Egyptian, Gulf, Levantine, North African, and MSA, with an accuracy of 52%. We discuss dialect identification errors in the context of dialect code-switching between Dialectal Arabic and MSA, and compare the error pattern between manually labeled data, and the output from our classifier. We also release the train and test data as standard corpus for dialect identification.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Spoken language identification | Untranscribed mixed-speech dataset | SVM | ACC | 45.2% | #1 of 4 | Archive leaderboard | report |
| Spoken language identification | Untranscribed mixed-speech dataset | SVM | PRC | 44.8% | #1 of 4 | Archive leaderboard | report |
| Spoken language identification | Untranscribed mixed-speech dataset | SVM | RCL | 45.4% | #1 of 4 | Archive leaderboard | report |
| Spoken language identification | Untranscribed mixed-speech dataset | n-gram Language Model | ACC | 40.4% | #2 of 4 | Archive leaderboard | report |
| Spoken language identification | Untranscribed mixed-speech dataset | n-gram Language Model | PRC | 40.2% | #2 of 4 | Archive leaderboard | report |
| Spoken language identification | Untranscribed mixed-speech dataset | n-gram Language Model | RCL | 41.3% | #2 of 4 | Archive leaderboard | report |
| Spoken language identification | Untranscribed mixed-speech dataset | Max Ent | ACC | 40% | #3 of 4 | Archive leaderboard | report |
| Spoken language identification | Untranscribed mixed-speech dataset | Max Ent | PRC | 40% | #3 of 4 | Archive leaderboard | report |
| Spoken language identification | Untranscribed mixed-speech dataset | Max Ent | RCL | 40.6% | #3 of 4 | Archive leaderboard | report |
| Spoken language identification | Untranscribed mixed-speech dataset | Naive Bayes | ACC | 37.9% | #4 of 4 | Archive leaderboard | report |
| Spoken language identification | Untranscribed mixed-speech dataset | Naive Bayes | PRC | 37.5% | #4 of 4 | Archive leaderboard | report |
| Spoken language identification | Untranscribed mixed-speech dataset | Naive Bayes | RCL | 50.2% | #4 of 4 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections