Browse State-of-the-Art › Dialect Identification
Dialect Identification
33 papers with code · 0 benchmarks · 3 datasets archive 2025-07-28
Dialectal Arabic Identification
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 33 papers with code (189 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
24 Oct 2023 3 repositories listedSeveral recent papers have published good solutions for language identification (LID) for about 300 high-resource and medium-resource languages.
-
23 Dec 2024 1 repository listedThis paper presents a novel approach to fine-tuning the Qwen2-1.
-
4 Oct 2024 1 repository listedTo provide benchmarks and simultaneously demonstrate the challenges of our dataset, we fine-tune state-of-the-art pre-trained models for two downstream tasks: (1) Dialect identification and (2) Speech recognition.
-
19 Mar 2024 1 repository listedNamed Entity Recognition (NER) is a fundamental task to extract key information from texts, but annotated resources are scarce for dialects.
-
25 Oct 2023 1 repository listedWe present ArTST, a pre-trained Arabic text and speech transformer for supporting open-source speech technologies for the Arabic language.
-
20 Oct 2023 1 repository listedAutomatic Arabic Dialect Identification (ADI) of text has gained great popularity since it was introduced in the early 2010s.
-
20 Oct 2023 1 repository listedTranscribed speech and user-generated text in Arabic typically contain a mixture of Modern Standard Arabic (MSA), the standardized language taught in schools, and Dialectal Arabic (DA), used in daily communications.
-
22 May 2023 1 repository listedWe show that DADA is effective for both single task and instruction finetuned language models, offering an extensible and interpretable framework for adapting existing LLMs to different English dialects.
-
19 May 2023 1 repository listedThe North S\'{a}mi (NS) language encapsulates four primary dialectal variants that are related but that also have differences in their phonology, morphology, and vocabulary.
-
18 May 2023 1 repository listedIn this work, we explore Parameter-Efficient-Learning (PEL) techniques to repurpose a General-Purpose-Speech (GSM) model for Arabic dialect identification (ADI).
-
6 Mar 2023 1 repository listedOur proposed approach consists of a two-stage system and outperforms other participants' systems and previous works in this domain.
-
15 Dec 2022 1 repository listedWe present a novel corpus for French dialect identification comprising 413, 522 French text samples collected from public news websites in Belgium, Canada, France and Switzerland.
-
22 Oct 2022 1 repository listedContrastive learning (CL) brought significant progress to various NLP tasks.
-
18 Oct 2022 1 repository listedWe describe findings of the third Nuanced Arabic Dialect Identification Shared Task (NADI 2022).
-
1 Oct 2022 1 repository listedThis report presents the results of the shared tasks organized as part of the VarDial Evaluation Campaign 2022.
-
1 Oct 2022 1 repository listedThis article describes the language identification approach used by the SUKI team in the Identification of Languages and Dialects of Italy and the French Cross-Domain Dialect Identification shared tasks organized as…
-
23 Dec 2021 1 repository listedIn this work, we introduce three light and fast versions of distilled BERT models for the Romanian language: Distil-BERT-base-ro, Distil-RoBERT-base, and DistilMulti-BERT-base-ro.
-
6 Nov 2021 1 repository listedFinnish is a language with multiple dialects that not only differ from each other in terms of accent (pronunciation) but also in terms of morphological forms and lexical choice.
-
2 Aug 2021 1 repository listedTo address this issue, we propose a new architecture, named dynamic multi-scale convolution, which consists of dynamic kernel convolution, local multi-scale learning, and global multi-scale pooling.
-
7 May 2021 1 repository listedThis paper releases "AraCOVID19-MFH" a manually annotated multi-label Arabic COVID-19 fake news and hate speech detection dataset.
-
1 Apr 2021 1 repository listedIn this work we compare the performance of convolutional neural networks and shallow models on three out of the four language identification shared tasks proposed in the VarDial Evaluation Campaign 2021.
-
4 Mar 2021 1 repository listedThis Shared Task includes four subtasks: country-level Modern Standard Arabic (MSA) identification (Subtask 1.
-
Adapting MARBERT for Improved Arabic Dialect Identification: Submission to the NADI 2021 Shared Task1 Mar 2021 1 repository listedTasks are to identify the geographic origin of short Dialectal (DA) and Modern Standard Arabic (MSA) utterances at the levels of both country and province.
-
1 Dec 2020 1 repository listedIn the last few years, deep learning has proved to be a very effective paradigm to discover patterns in large data sets.
-
1 Dec 2020 1 repository listedIn this paper we present the architecture, processing pipeline and results of the ensemble model developed for Romanian Dialect Identification task.
-
10 Oct 2020 1 repository listedAlthough the prediction of dialects is an important language processing task, with a wide range of applications, existing work is largely limited to coarse-grained varieties.
-
30 Jul 2020 1 repository listedWe conduct a subjective evaluation by human annotators, showing that humans attain much lower accuracy rates compared to machine learning (ML) models.
-
10 Jul 2020 1 repository listedOur winning solution itself came in the form of an ensemble of different training iterations of our pre-trained BERT model, which achieved a micro-averaged F1-score of 26.
-
1 May 2020 1 repository listedWe present CAMeL Tools, a collection of open-source tools for Arabic natural language processing in Python.
-
21 Sep 2017 1 repository listedTwo hours of audio per dialect were released for development and a further two hours were used for evaluation.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections