Browse State-of-the-Art › Language Identification
Language Identification
143 papers with code · 6 benchmarks · 20 datasets archive 2025-07-28
Language identification is the task of determining the language of a text.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| VOXLINGUA107 (2 rows) | XLS-R | XLS-R: Self-supervised Cross-lingual Speech Representation... | code | — | Compare |
| GlotLID-C (1 row) | GlotLID | GlotLID: Language Identification for Low-Resource Languages | code | — | Compare |
| Nordic Language Identification (1 row) | FastText | Discriminating Between Similar Nordic Languages | code | — | Compare |
| OpenSubtitles (1 row) | Apple bi-LSTM | A reproduction of Apple's bi-directional LSTM models for language... | code | — | Compare |
| Universal Dependencies (1 row) | Apple bi-LSTM | A reproduction of Apple's bi-directional LSTM models for language... | code | — | Compare |
| VoxForge (1 row) | ConformerG-P | BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised... | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
20 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 143 papers with code (794 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
22 May 2023 4 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedExpanding the language coverage of speech technology has the potential to improve access to information for many more people.
-
8 Jun 2021 4 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedSpeechBrain is an open-source and all-in-one speech toolkit.
-
23 Jan 2018 4 repositories listedThis paper describes the WiLI-2018 benchmark dataset for monolingual written natural language identification.
-
24 Oct 2023 3 repositories listedSeveral recent papers have published good solutions for language identification (LID) for about 300 high-resource and medium-resource languages.
-
23 May 2025 2 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedDespite these advancements, CosyVoice 2 exhibits limitations in language coverage, domain diversity, data volume, text formats, and post-training techniques.
-
31 Oct 2024 2 repositories listedThe need for large text corpora has increased with the advent of pretrained language models and, in particular, the discovery of scaling laws for these models.
-
11 Oct 2024 2 repositories listedIn this paper, we investigate the use of N-gram models and Large Pre-trained Multilingual models for Language Identification (LID) across 11 South African languages.
-
21 Feb 2024 2 repositories listedSemEval-2024 Task 8 is focused on multigenerator, multidomain, and multilingual black-box machine-generated text detection.
-
17 Nov 2021 2 repositories listedOn the CoVoST-2 speech translation benchmark, we improve the previous state of the art by an average of 7.
-
11 Dec 2020 2 repositories listedAutomatic language identification is a challenging problem.
-
1 Dec 2020 2 repositories listedThis paper describes the systems our team (AdelaideCyC) has developed for SemEval Task 12 (OffensEval 2020) to detect offensive language in social media.
-
1 Dec 2020 2 repositories listedThis task has been organized for several languages, e.
-
25 Nov 2020 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Speech activity detection and speaker diarization are used to extract segments from the videos that contain speech.
-
13 Dec 2019 2 repositories listedTo our knowledge this is the largest audio corpus in the public domain for speech recognition, both in terms of number of hours and number of languages.
-
22 Oct 2019 2 repositories listedRecent breakthroughs in deep learning often rely on representation learning and knowledge transfer.
-
19 Mar 2019 2 repositories listed Syntology ran 0 of 10 samples · 10 unverifiedWe present the results and the main findings of SemEval-2019 Task 6 on Identifying and Categorizing Offensive Language in Social Media (OffensEval).
-
16 Apr 2018 2 repositories listedWe present a treebank of Hindi-English code-switching tweets under Universal Dependencies scheme and propose a neural stacking model for parsing that efficiently leverages part-of-speech tag and syntactic tree…
-
10 May 2025 1 repository listedThe digital exclusion of endangered languages remains a critical challenge in NLP, limiting both linguistic research and revitalization efforts.
-
9 Mar 2025 1 repository listedAutomatic language identification is frequently framed as a multi-class classification problem.
-
20 Feb 2025 1 repository listedIn this study, we conduct the first comprehensive evaluation of machine translation (MT) performance on bug reports, analyzing the capabilities of DeepL, AWS Translate, and large language models such as ChatGPT, Claude,…
-
10 Feb 2025 1 repository listedIdentifying closely related languages at sentence level is difficult, in particular because it is often impossible to assign a sentence to a single language.
-
27 Jan 2025 1 repository listedEndangered languages, such as Navajo - the most widely spoken Native American language - are significantly underrepresented in contemporary language technologies, exacerbating the challenges of their preservation and…
-
10 Jan 2025 1 repository listedWhile recent multilingual automatic speech recognition models claim to support thousands of languages, ASR for low-resource languages remains highly unreliable due to limited bimodal speech and text training data.
-
30 Sep 2024 1 repository listedIn this work, we present AfriHuBERT, an extension of mHuBERT-147, a compact self-supervised learning (SSL) model pretrained on 147 languages.
-
27 Sep 2024 1 repository listedMultilingual Automatic Speech Recognition (ASR) models are typically evaluated in a setting where the ground-truth language of the speech utterance is known, however, this is often not the case for most practical…
-
11 Aug 2024 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedBeam search decoding is the de-facto method for decoding auto-regressive Neural Machine Translation (NMT) models, including multilingual NMT where the target language is specified as an input.
-
7 Aug 2024 1 repository listedWe present Speech-MASSIVE, a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus.
-
25 Jun 2024 1 repository listedLanguage identification is used as the first step in many data collection and crawling efforts because it allows us to sort online text into language-specific buckets.
-
13 Jun 2024 1 repository listedSelf-supervised learning (SSL) representations from massively multilingual models offer a promising solution for low-resource language speech tasks.
-
10 Jun 2024 1 repository listedThis method uses the LID itself to identify the features that require masking and does not rely on any external resource.
Syntology lines on 6 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections