Home › Datasets › task › Language Identification
Language Identification datasets
archive 2025-07-28
20 datasets carry the task tag "Language Identification" (the task itself: Language Identification), ordered by the archive's paper count. Page 1 of 1: 20 shown of 20. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Language Identification datasets 1–20 of 20
The Universal Dependencies (UD) project seeks to develop cross-linguistically consistent treebank annotation of morphology and syntax for multiple languages.
520 papers · 5 benchmarks
Common Voice is an audio dataset that consists of a unique MP3 and corresponding text file.
449 papers · 143 benchmarks
OpenSubtitles is collection of multilingual parallel corpora.
214 papers · 3 benchmarks
OLID (Offensive Language Identification Dataset)
The OLID is a hierarchical dataset to identify the type and the target of offensive texts in social media.
152 papers · 1 benchmark
CONAN (COunter NArratives through Nichesourcing)
COunter NArratives through Nichesourcing (CONAN) is a dataset that consists of 4,078 pairs over the 3 languages.
27 papers · 0 benchmarks
MOROCO (MOldavian and ROmanian Dialectal COrpus)
The MOldavian and ROmanian Dialectal COrpus (MOROCO) is a corpus that contains 33,564 samples of text (with over 10 million tokens) collected from the news domain.
22 papers · 0 benchmarks
The Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian languages.
14 papers · 0 benchmarks
VoxForge is an open speech dataset that was set up to collect transcribed speech for use with Free and Open Source Speech Recognition Engines (on Linux, Windows and Mac).
11 papers · 9 benchmarks
A parallel corpus of Hindi and English, and HindMonoCorp, a monolingual corpus of Hindi in their release version 0.5.
5 papers · 0 benchmarks
OGTD (Offensive Greek Tweet Dataset)
A manually annotated dataset containing 4,779 posts from Twitter annotated as offensive and not offensive.
3 papers · 0 benchmarks
WiLI-2018 is a benchmark dataset for monolingual written natural language identification.
3 papers · 0 benchmarks
Language Identification Dataset
2 papers · 2 benchmarks
Clean version of UDHR (Universal Declaration of Human Rights), at the long sentence level.
2 papers · 0 benchmarks
Collection of news websites in low-resource languages.
1 paper · 0 benchmarks
StoryBooks for 174 unique languages.
1 paper · 0 benchmarks
L3Cube-MahaCorpus is a Marathi monolingual data set scraped from different internet sources.
1 paper · 0 benchmarks
The first Portuguese dataset compiled for Native Language Identification (NLI), the task of identifying an author's first language based on their second language writing.
1 paper · 0 benchmarks
Automatic language identification is a challenging problem.
1 paper · 1 benchmark
TuGebic (A Turkish-German Bilingual Code-Switching Corpus)
TuGebic is a corpus of recordings of spontaneous speech samples from Turkish-German bilinguals, and the compilation of a corpus called TuGebic.
1 paper · 0 benchmarks
The English-Pashto Language Dataset (EPLD) is a comprehensive resource aimed to provide linguistic insights into the Pashto language.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.