Home › Datasets › language › Turkish
Turkish datasets
archive 2025-07-28
63 datasets carry the language tag "Turkish", ordered by the archive's paper count. Page 1 of 2: 48 shown of 63. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 language tags shown of 367, by dataset count; the full filter by modality, task and language is on /datasets
Turkish datasets 1–48 of 63
The Universal Dependencies (UD) project seeks to develop cross-linguistically consistent treebank annotation of morphology and syntax for multiple languages.
520 papers · 5 benchmarks
The MPII Human Pose Dataset for single person pose estimation is composed of about 25K images of which 15K are training samples, 3K are validation samples and 7K are testing samples (which labels are withheld by the authors).
495 papers · 4 benchmarks
Common Voice is an audio dataset that consists of a unique MP3 and corresponding text file.
449 papers · 143 benchmarks
XNLI (Cross-lingual Natural Language Inference)
The Cross-lingual Natural Language Inference (XNLI) corpus is the extension of the Multi-Genre NLI (MultiNLI) corpus to 15 languages.
349 papers · 7 benchmarks
MuST-C currently represents the largest publicly available multilingual corpus (one-to-many) for speech translation.
216 papers · 2 benchmarks
XQuAD (Cross-lingual Question Answering Dataset) is a benchmark dataset for evaluating cross-lingual question answering performance.
190 papers · 1 benchmark
This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages.
115 papers · 0 benchmarks
WikiANN, also known as PAN-X, is a multilingual named entity recognition dataset.
72 papers · 4 benchmarks
CoVoST2 (Common Voice Speech-To-Text 2)
End-to-end speech-to-text translation (ST) has recently witnessed an increased interest given its system simplicity, lower inference latency and less compounding errors compared to cascaded ST (i.e.
66 papers · 0 benchmarks
OSCAR or Open Super-large Crawled ALMAnaCH coRpus is a huge multilingual corpus obtained by language classification and filtering of the Common Crawl corpus using the goclassy architecture.
64 papers · 0 benchmarks
XL-Sum is a comprehensive and diverse dataset for abstractive summarization comprising 1 million professionally annotated article-summary pairs from BBC, extracted using a set of carefully designed heuristics.
64 papers · 0 benchmarks
WikiLingua includes ~770k article and summary pairs in 18 languages from WikiHow.
52 papers · 1 benchmark
MKQA (Multilingual Knowledge Questions and Answers)
Multilingual Knowledge Questions and Answers (MKQA) is an open-domain question answering evaluation set comprising 10k question-answer pairs aligned across 26 typologically diverse languages (260k question-answer pairs in total).
49 papers · 0 benchmarks
MLSUM (MultiLingual SUMmarization)
A large-scale MultiLingual SUMmarization dataset.
45 papers · 4 benchmarks
CoVoST is a large-scale multilingual speech-to-text translation corpus.
34 papers · 0 benchmarks
IGLUE (Image-Grounded Language Understanding Evaluation)
The Image-Grounded Language Understanding Evaluation (IGLUE) benchmark brings together—by both aggregating pre-existing datasets and creating new ones—visual question answering, cross-modal retrieval, grounded reasoning, and grounded…
31 papers · 0 benchmarks
MaRVL (Multicultural Reasoning over Vision and Language)
Multicultural Reasoning over Vision and Language (MaRVL) is a dataset based on an ImageNet-style hierarchy representative of many languages and cultures (Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish).
29 papers · 1 benchmark
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets.
28 papers · 0 benchmarks
CVSS is a massively multilingual-to-English speech to speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English.
26 papers · 1 benchmark
OASST1 (OpenAssistant Conversations Dataset)
license: apache-2.0 tags: human-feedback sizecategories: 100K Languages with under 1000 messages Vietnamese: 952 Basque: 947 Polish: 886 Hungarian: 811 Arabic: 666 Dutch: 628 Swedish: 512 Turkish: 454 Finnish: 386 Czech: 372 Danish: 358…
25 papers · 0 benchmarks
News translation is a recurring WMT task.
24 papers · 0 benchmarks
XGLUE is an evaluation benchmark XGLUE,which is composed of 11 tasks that span 19 languages.
22 papers · 2 benchmarks
xSID (Cross-lingual Slot and Intent Detection)
xSID, a new evaluation benchmark for cross-lingual (X) Slot and Intent Detection in 13 languages from 6 language families, including a very low-resource dialect, covering Arabic (ar), Chinese (zh), Danish (da), Dutch (nl), English (en),…
18 papers · 0 benchmarks
X-FACT is a large publicly available multilingual dataset for factual verification of naturally existing real-world claims.
16 papers · 0 benchmarks
News translation is a recurring WMT task.
8 papers · 0 benchmarks
We introduce HumanEval-XL, a massively multilingual code generation benchmark specifically crafted to address this deficiency.
6 papers · 0 benchmarks
RELX is a benchmark dataset for cross-lingual relation classification in English, French, German, Spanish and Turkish.
6 papers · 0 benchmarks
XL-BEL is a benchmark for cross-lingual biomedical entity linking (XL-BEL).
6 papers · 0 benchmarks
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
In this paper, we present AnlamVer, which is a semantic model evaluation dataset for Turkish designed to evaluate word similarity and word relatedness tasks while discriminating those two relations from each other.
4 papers · 0 benchmarks
DISRPT2019 (DISRPT2019 shared task on Discourse Unit Segmentation and Connective Detection)
The DISRPT 2019 workshop introduces the first iteration of a cross-formalism shared task on discourse unit segmentation.
4 papers · 0 benchmarks
GeoCoV19 is a large-scale Twitter dataset containing more than 524 million multilingual tweets.
4 papers · 0 benchmarks
DISRPT2021 (DISRPT2021 shared task on Discourse Unit Segmentation, Connective Detection and Discourse Relation Classification)
The DISRPT 2021 shared task, co-located with CODI 2021 at EMNLP, introduces the second iteration of a cross-formalism shared task on discourse unit segmentation and connective detection, as well as the first iteration of a cross-formalism…
3 papers · 0 benchmarks
DivEMT (Post-Editing Effort Across Typologically-diverse Languages)
DivEMT, the first publicly available post-editing study of Neural Machine Translation (NMT) over a typologically diverse set of target languages.
3 papers · 0 benchmarks
xMIND (A Multilingual Dataset for Cross-lingual News Recommendation)
xMIND is an open, large-scale multilingual news dataset for multi- and cross-lingual news recommendation.
3 papers · 0 benchmarks
AM2iCo (Adversarial and Multilingual Meaning in Context)
AM2iCo is a wide-coverage and carefully designed cross-lingual and multilingual evaluation set.
2 papers · 0 benchmarks
The fetoscopy placenta dataset is associated with our MICCAI2020 publication titled “Deep Placental Vessel Segmentation for Fetoscopic Mosaicking”.
2 papers · 0 benchmarks
M2QA (Multi-domain Multilingual Question Answering)
M2QA (Multi-domain Multilingual Question Answering) is an extractive question answering benchmark for evaluating joint language and domain transfer.
2 papers · 0 benchmarks
MTST (Mobile Turkish Scene Text)
The Mobile Turkish Scene Text (MTST 200) dataset consists of 200 indoor and outdoor Turkish scene text images.
2 papers · 0 benchmarks
MultiTACRED is a multilingual version of the large-scale TAC Relation Extraction Dataset.
2 papers · 0 benchmarks
The dataset contains training and evaluation data for 12 languages: - Vietnamese - Romanian - Latvian - Czech - Polish - Slovak - Irish - Hungarian - French - Turkish - Spanish - Croatian For each language, one training, one development…
2 papers · 12 benchmarks
NLI-TR (Natural Language Inference in Turkish)
Natural Language Inference in Turkish (NLI-TR) provides translations of two large English NLI datasets into Turkish and had a team of experts validate their translation quality and fidelity to the original labels.
2 papers · 0 benchmarks
Bianet is a parallel news corpus in Turkish, Kurdish and English It contains 3,214 Turkish articles with their sentence-aligned Kurdish or English translations from the Bianet online newspaper.
1 paper · 0 benchmarks
GLAMI-1M (A Multilingual Image-Text Fashion Dataset)
We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark.
1 paper · 1 benchmark
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
1 paper · 0 benchmarks
Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.
1 paper · 0 benchmarks
LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali,…
1 paper · 0 benchmarks
Describe the Marmara Turkish Coreference Corpus, which is an annotation of the whole METU-Sabanci Turkish Treebank with mentions and coreference chains.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.