Papers › Automatic Dialect Detection in Arabic Broadcast Speech

Automatic Dialect Detection in Arabic Broadcast Speech

23 Sep 2015arXiv:1509.06928archive 2025-07-28

Ahmed Ali, Najim Dehak, Patrick Cardinal, Sameer Khurana, Sree Harsha Yella, James Glass, Peter Bell, Steve Renals

We investigate different approaches for dialect identification in Arabic broadcast speech, using phonetic, lexical features obtained from a speech recognition system, and acoustic features using the i-vector framework. We studied both generative and discriminate classifiers, and we combined these features using a multi-class Support Vector Machine (SVM). We validated our results on an Arabic/English language identification task, with an accuracy of 100%. We used these features in a binary classifier to discriminate between Modern Standard Arabic (MSA) and Dialectal Arabic, with an accuracy of 100%. We further report results using the proposed method to discriminate between the five most widely used dialects of Arabic: namely Egyptian, Gulf, Levantine, North African, and MSA, with an accuracy of 52%. We discuss dialect identification errors in the context of dialect code-switching between Dialectal Arabic and MSA, and compare the error pattern between manually labeled data, and the output from our classifier. We also release the train and test data as standard corpus for dialect identification.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Dialect IdentificationLanguage IdentificationSpeech RecognitionSpoken language identificationspeech-recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Spoken language identification Untranscribed mixed-speech dataset SVM ACC 45.2% #1 of 4 Archive leaderboard report
Spoken language identification Untranscribed mixed-speech dataset SVM PRC 44.8% #1 of 4 Archive leaderboard report
Spoken language identification Untranscribed mixed-speech dataset SVM RCL 45.4% #1 of 4 Archive leaderboard report
Spoken language identification Untranscribed mixed-speech dataset n-gram Language Model ACC 40.4% #2 of 4 Archive leaderboard report
Spoken language identification Untranscribed mixed-speech dataset n-gram Language Model PRC 40.2% #2 of 4 Archive leaderboard report
Spoken language identification Untranscribed mixed-speech dataset n-gram Language Model RCL 41.3% #2 of 4 Archive leaderboard report
Spoken language identification Untranscribed mixed-speech dataset Max Ent ACC 40% #3 of 4 Archive leaderboard report
Spoken language identification Untranscribed mixed-speech dataset Max Ent PRC 40% #3 of 4 Archive leaderboard report
Spoken language identification Untranscribed mixed-speech dataset Max Ent RCL 40.6% #3 of 4 Archive leaderboard report
Spoken language identification Untranscribed mixed-speech dataset Naive Bayes ACC 37.9% #4 of 4 Archive leaderboard report
Spoken language identification Untranscribed mixed-speech dataset Naive Bayes PRC 37.5% #4 of 4 Archive leaderboard report
Spoken language identification Untranscribed mixed-speech dataset Naive Bayes RCL 50.2% #4 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections