{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/automatic-dialect-detection-in-arabic","title":"Automatic Dialect Detection in Arabic Broadcast Speech","arxiv_id":"1509.06928","date":"2015-09-23","proceeding":null,"authors":["Ahmed Ali","Najim Dehak","Patrick Cardinal","Sameer Khurana","Sree Harsha Yella","James Glass","Peter Bell","Steve Renals"],"abstract":"We investigate different approaches for dialect identification in Arabic\nbroadcast speech, using phonetic, lexical features obtained from a speech\nrecognition system, and acoustic features using the i-vector framework. We\nstudied both generative and discriminate classifiers, and we combined these\nfeatures using a multi-class Support Vector Machine (SVM). We validated our\nresults on an Arabic/English language identification task, with an accuracy of\n100%. We used these features in a binary classifier to discriminate between\nModern Standard Arabic (MSA) and Dialectal Arabic, with an accuracy of 100%. We\nfurther report results using the proposed method to discriminate between the\nfive most widely used dialects of Arabic: namely Egyptian, Gulf, Levantine,\nNorth African, and MSA, with an accuracy of 52%. We discuss dialect\nidentification errors in the context of dialect code-switching between\nDialectal Arabic and MSA, and compare the error pattern between manually\nlabeled data, and the output from our classifier. We also release the train and\ntest data as standard corpus for dialect identification.","url_abs":"http://arxiv.org/abs/1509.06928v2","url_pdf":"http://arxiv.org/pdf/1509.06928v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"automatic-dialect-detection-in-arabic","repo_url":"https://github.com/Qatar-Computing-Research-Institute/dialectID","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"dialect-identification","task_name":"Dialect Identification"},{"task_slug":"language-identification","task_name":"Language Identification"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"spoken-language-identification","task_name":"Spoken language identification"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/spoken-language-identification-on-1","task":"Spoken language identification","dataset":"Untranscribed mixed-speech dataset","model":"SVM","rank_in_archive_order":1,"of":4,"metrics":{"ACC":"45.2%","PRC":"44.8%","RCL":"45.4%"},"uses_additional_data":false},{"leaderboard":"/sota/spoken-language-identification-on-1","task":"Spoken language identification","dataset":"Untranscribed mixed-speech dataset","model":"n-gram Language Model","rank_in_archive_order":2,"of":4,"metrics":{"ACC":"40.4%","PRC":"40.2%","RCL":"41.3%"},"uses_additional_data":false},{"leaderboard":"/sota/spoken-language-identification-on-1","task":"Spoken language identification","dataset":"Untranscribed mixed-speech dataset","model":"Max Ent","rank_in_archive_order":3,"of":4,"metrics":{"ACC":"40%","PRC":"40%","RCL":"40.6%"},"uses_additional_data":false},{"leaderboard":"/sota/spoken-language-identification-on-1","task":"Spoken language identification","dataset":"Untranscribed mixed-speech dataset","model":"Naive Bayes","rank_in_archive_order":4,"of":4,"metrics":{"ACC":"37.9%","PRC":"37.5%","RCL":"50.2%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1509.06928","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}